Stroke-Level Connectivity Verification: Grounding Vision-Language Models Against Topology Hallucination in Diagram Understanding
Md Ashikur Rahman is a Lead AI Engineer at The KOW Company, specializing in multimodal AI, vision-language systems, generative AI, and applied computer vision.
Research interests: trustworthy multimodal learning; vision-language models; AI safety and robustness.
He currently leads a multidisciplinary team of 15+ engineers and researchers and delivers enterprise AI products used by 200+ brands. He architected computer-vision systems that have processed more than 4.5 million images globally.
Recent Highlights
- ICDAR 2026 oral paper on diagram understanding with LogicBench-1K (research lead and corresponding author).
- ICCA 2026 paper on uncertainty-calibrated retrieval for medical vision-language hallucination reduction (last author).
- Manuscripts under review on conformal risk control for LLM tool calls, detector-based grounding metrics in video QA, math-encoded jailbreaks, and residual stream rebalancing.
- The Fitting Room: cross-brand virtual try-on platform supporting products from 170+ brands.
- CogniX: early-stage R&D for proprietary multimodal content creation.
Education
B.Sc. in Computer Science and Engineering, American International University-Bangladesh, 2011-2015. CGPA 3.87/4.00 (WES equivalent 3.94/4.00); magna cum laude; top 3%; Merit Scholarship and Tuition Fee Waiver.
Publications
UCAR: Uncertainty-Calibrated Adaptive Retrieval for Hallucination Reduction in Medical Vision-Language Models
Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls
When Detector-Based Grounding Metrics Measure Vocabulary: A Cautionary Audit of Entity Claims in Video-QA Reasoning Traces
Decoding Harm: Do Reasoning Models Resist Math-Encoded Jailbreaks?
Residual Stream Rebalancing: Training-Free Hallucination Mitigation in Vision-Language Models
Automated Detection of Diabetic Retinopathy Using Deep Residual Learning
Selected Projects
- Developed salient-object segmentation for background removal; improved segmentation quality by 17% on internal benchmarks, reduced processing time by 30%, enabled uploads of up to 257 MB, and achieved a 2.27-second average processing time across standard workloads.
- Scaled Retouched.ai to process 4.5M+ images globally for hundreds of customers using PyTorch, U²-Net-inspired salient-object segmentation, FastAPI, and Google Cloud Platform.
- Co-designed and launched image- and video-generation APIs used by 200+ brands for creative and product-image workflows.
- Engineered workflows for reference-image conditioning, asynchronous processing, prompt classification, intent routing, and automated model selection.
- Conceived and co-led a unified cross-brand virtual try-on platform supporting products from 170+ brands.
- Architected a Dockerized FastAPI/Nginx backend using SQL Server, Google Cloud Storage, Redis, recommendation services, and 2D virtual try-on pipelines.
- Leading the development of an AI catalog-audit and image quality assurance platform for a major US retail client.
- Automating image QA, catalog validation, metadata checks, and content-health monitoring across DAM and CMS workflows using NLP, PyTorch, and FastAPI.
- Leading early-stage R&D for multimodal product-visualization and editing workflows.
- Researching and prototyping multi-reference fusion and editing workflows for garment replacement, scene composition, product visualization, and style transfer using PyTorch, Diffusers, ComfyUI, and parameter-efficient fine-tuning (LoRA, QLoRA).
Experience
- Lead a multidisciplinary team of 15+ ML engineers, software engineers, and junior researchers, driving technical strategy, system architecture, engineering execution, quality standards, and cross-functional delivery.
- Lead the development and deployment of scalable AI systems spanning virtual try-on, computer vision, 3D reconstruction, and multimodal content creation (Retouched.ai, Omnimage.ai, HoloSnap.ai, CogniX, and The Fitting Room).
- Direct applied research on AI hallucination, prompt safety, visual grounding, and video understanding, operationalizing research into production-ready models and evaluation frameworks.
- Mentor and supervise four junior researchers in literature review, research planning, experimental design, reproducible implementation, result analysis, and academic writing.
- Develop and deploy deep learning-based computer vision and image-processing systems; publish reproducible code, datasets, and model checkpoints on GitHub and Hugging Face.
- Improved object detection and segmentation performance by 20-35% across internal evaluation benchmarks; the resulting models were later deployed in Retouched.ai.
- Led 6+ client ML engagements from requirements gathering through production delivery; built offline evaluation pipelines and A/B testing workflows to validate model quality, inference performance, and business and operational outcomes.
- Built deep learning models for object recognition, image segmentation, and background-removal workflows; developed scalable preprocessing, training, and A/B testing pipelines.
- Led .NET and SQL Server supply-chain systems on a 4 TB database, reducing report generation from 20 minutes to 40-54 seconds.
- Achieved 70-75% process automation and 99.9% synchronization success for offline-capable enterprise workflows.
- Developed ASP.NET MVC features, backend services, and database integrations for DevSkill.com.
Technical Skills
- Multimodal and Vision-Language AI
- Python · PyTorch · Vision-Language Models · Visual Grounding · Hallucination Evaluation · Multimodal Learning
- Generative AI and LLMs
- Diffusion Models · Large Language Models (including Llama and Qwen) · Prompt Safety · Parameter-Efficient Fine-Tuning (LoRA, QLoRA)
- Computer Vision
- Object Detection · Segmentation · Pose Estimation · Image Quality Assurance
- 3D Vision
- Structure from Motion · Multi-View Stereo · COLMAP · Open3D · Neural Radiance Fields (NeRF) · 3D Gaussian Splatting
- Production ML and MLOps
- Model Serving · Deployment · Offline and Online Evaluation · A/B Testing · Dataset Design
- Backend and Cloud
- FastAPI · Docker · Nginx · Redis · Google Cloud Platform · Google Cloud Storage · SQL Server
Honors and Awards
- Champion, BASIS National ICT Awards, 2020 (Retouched.ai)
- Finalist, Asia Pacific ICT Alliance Awards, 2021
- Artificial Intelligence in Advertising, invited workshop speaker, Daffodil International University
- magna cum laude; Top 3%; Merit Scholarship and Tuition Fee Waiver, AIUB
Additional Academic Training
Individual Study in Mathematics for Research, 2025 (through The KOW Company): Multivariable Calculus with Dr. Anindita Paul; Differential Geometry with Dr. A. K. M. Nazimuddin, Department of Mathematics and Data Science, East West University.
Contact
Open to research collaboration, speaking invitations, and professional inquiries in multimodal AI, vision-language models, and applied machine learning.
- mdashikur.rafi@gmail.com
- GitHub
- Google Scholar
- Dhaka, Bangladesh