Imtiaz Mashrafee

Résumé

Imtiaz Mashrafee Samin

AI Engineer

AI Engineer with industry experience building and deploying LLM, RAG, and machine-learning systems. Experienced across PyTorch, retrieval systems, Docker/CI/CD, and production AI deployments, with research focused on transformer robustness and backdoor defense.

01 Work experience

AI Engineer at Data Solution-360

Jan 2026 to PresentDhaka, Bangladesh

Worked on Linx360.com, an AI-powered interview and assessment platform that identifies skill gaps, delivers role-aligned interview practice, and helps candidates become job-ready.

  • Built a Python/FastAPI MCQ assessment pipeline with OpenAI API, generating 15 personalized questions from candidate profiles and skill gaps; parallelized requests with Redis TTL storage, cutting latency from about 10 to 15 s down to about 2 to 5 s.
  • Built an LLM critic/regeneration loop with strict JSON schemas and Bloom-aligned rubrics, regenerating rejected questions until all 15 passed validation.
  • Engineered a scikit-learn ranking/scoring workflow over about 30K records using Pandas/NumPy for feature engineering, leakage-safe proxy targets, and held-out evaluation with Precision@50.
  • Migrated Docker services from Linux to DigitalOcean App Platform with automated GitHub Actions and GHCR CI/CD.
  • Used Claude Code and Codex with MCP, Superpowers, and pytest TDD to plan, implement, debug, and validate features end to end.
  • Python
  • FastAPI
  • OpenAI API
  • Pandas
  • NumPy
  • scikit-learn
  • Redis
  • Docker
  • GitHub Actions
  • GHCR
  • DigitalOcean
  • pytest
  • Claude Code
  • OpenAI Codex
  • MCP
  • Superpowers

02 Projects

Data-Aware RAG System for Research Papers

Local multi-paper research assistant with hybrid retrieval and source/page citations.

  • Built Python PDF ingestion with Marker and pdfplumber, using page-aware chunking to preserve source and page metadata.
  • Generated local all-MiniLM-L6-v2 embeddings with Sentence Transformers and combined dense cosine search with BM25 using normalized weighted fusion.
  • Decomposed multi-paper questions into atomic needs, then rewrote and routed queries to retrieve evidence from relevant papers.
  • Generated citation-aware answers locally with Ollama, propagating retrieved paper and page metadata into responses.
  • Parallelized LLM query classification, reducing an eight-query experiment from 70.6 s to 20.3 s with the same 26/35 score.
  • Python
  • Sentence Transformers
  • all-MiniLM-L6-v2
  • BM25
  • NumPy
  • Marker
  • pdfplumber
  • SQLite
  • Streamlit
  • Ollama

ClientManager Pro: AI-Powered Client Management System

Full-stack client platform with project workflows, real-time communication, and AI-powered repository Q&A.

  • Built Next.js frontend and Express/Node.js REST APIs with MongoDB persistence for client and project workflows.
  • Implemented JWT authentication and role-based access control across multiple user roles and protected routes.
  • Built GitHub repository Q&A with Transformers.js and all-MiniLM-L6-v2, storing embeddings in Pinecone and answering retrieved-context questions with Gemini.
  • Reduced repeated AI computation with embedding/answer caching and query deduplication, avoiding unnecessary embedding and generation calls.
  • Implemented Socket.IO real-time chat alongside project workflows and automated document generation.
  • Next.js
  • Node.js
  • Express
  • MongoDB
  • JWT
  • Socket.IO
  • Pinecone
  • Transformers.js
  • all-MiniLM-L6-v2
  • Gemini

This is a summary of my CV. Figures such as latency and speed-up are as reported there; the case studies on this site say what was measured and what was not.