Stroke-Level Connectivity Verification: Grounding Vision-Language Models Against Topology Hallucination in Diagram Understanding
Md Ashikur Rahman is a Lead AI Engineer at The KOW Company, specializing in multimodal AI, vision-language systems, generative AI, and applied computer vision.
Research interests: trustworthy multimodal machine learning; vision-language models; AI safety and robustness.
He currently leads a multidisciplinary team of 15+ engineers and researchers and delivers enterprise AI products used by 200+ brands. He architected computer-vision systems that have processed more than 4.5 million images globally.
Recent Highlights
- ICDAR 2026 oral paper on diagram understanding with LogicBench-1K (last and corresponding author).
- ICCA 2026 paper on uncertainty-calibrated retrieval for medical vision-language hallucination reduction (last author).
- Manuscripts under review on conformal risk control for LLM tool calls, detector-based grounding metrics in video QA, math-encoded jailbreaks, and residual stream rebalancing.
- The Fitting Room: cross-brand virtual try-on platform supporting products from 170+ brands.
- CogniX: early-stage R&D for proprietary multimodal content creation.
Education
B.Sc. in Computer Science and Engineering, American International University-Bangladesh, 2011-2015. CGPA 3.87/4.00 (WES equivalent 3.94/4.00); magna cum laude; top 3%; Merit Scholarship and Tuition Fee Waiver.
Publications
UCAR: Uncertainty-Calibrated Adaptive Retrieval for Hallucination Reduction in Medical Vision-Language Models
Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls
When Detector-Based Grounding Metrics Measure Vocabulary: A Cautionary Audit of Entity Claims in Video-QA Reasoning Traces
Decoding Harm: Do Reasoning Models Resist Math-Encoded Jailbreaks?
Residual Stream Rebalancing: Training-Free Hallucination Mitigation in Vision-Language Models
Automated Detection of Diabetic Retinopathy Using Deep Residual Learning
Selected Projects
- Developed salient-object segmentation for background removal; improved segmentation quality by 17% on internal benchmarks, reduced processing time by 30%, enabled uploads of up to 257 MB, and achieved a 2.27-second average processing time across standard workloads.
- Scaled Retouched.ai to process 4.5M+ images globally for hundreds of customers using PyTorch, U²-Net-inspired salient-object segmentation, FastAPI, and Google Cloud Platform.
- Co-designed and launched image- and video-generation APIs used by 200+ brands for creative and product-image workflows.
- Engineered workflows for reference-image conditioning, asynchronous processing, prompt classification, intent routing, and automated model selection.
- Conceived and co-led a unified cross-brand virtual try-on platform supporting products from 170+ brands.
- Architected a Dockerized FastAPI/Nginx backend using SQL Server, Google Cloud Storage, Redis, recommendation services, and 2D virtual try-on pipelines.
- Leading the development of an AI catalog-audit and image quality assurance platform for a major US retail client.
- Automating image QA, catalog validation, metadata checks, and content-health monitoring across DAM and CMS workflows using NLP, PyTorch, and FastAPI.
- Leading early-stage R&D for multimodal product-visualization and editing workflows.
- Researching and prototyping multi-reference fusion and editing workflows for garment replacement, scene composition, product visualization, and style transfer using PyTorch, Diffusers, ComfyUI, and parameter-efficient fine-tuning (LoRA, QLoRA).
Experience
- Lead a multidisciplinary team of 15+ ML engineers, software engineers, and junior researchers, driving technical strategy, system architecture, engineering execution, quality standards, and cross-functional delivery.
- Lead the design, development, and deployment of scalable AI systems spanning virtual try-on, computer vision, 3D reconstruction, and multimodal content creation (Retouched.ai, Omnimage.ai, HoloSnap.ai, CogniX, and The Fitting Room).
- Direct applied research on AI hallucination, prompt safety, visual grounding, and video understanding, operationalizing research into production-ready models and evaluation frameworks.
- Mentor and supervise four researchers on literature review, research planning, experimental design, reproducible implementation, result analysis, and academic writing.
- Develop and deploy deep learning–based computer vision and image-processing systems; publish reproducible code, datasets, and model checkpoints on GitHub and Hugging Face.
- Improved object detection and segmentation performance by 20-35% across internal evaluation benchmarks; the resulting models were later deployed in Retouched.ai.
- Led 6+ client ML engagements from requirements gathering through production delivery; built offline evaluation pipelines and A/B testing workflows to validate model quality, inference performance, and business and operational outcomes.
- Built deep learning models for object recognition, image segmentation, and background-removal workflows; developed scalable preprocessing, training, and A/B testing pipelines.
- Led .NET and SQL Server supply-chain systems on a 4 TB database, reducing report generation from 20 minutes to 40-54 seconds.
- Achieved 70-75% process automation and 99.9% synchronization success for offline-capable enterprise workflows.
- Developed ASP.NET MVC features, backend services, and database integrations for DevSkill.com.
Technical Skills
- Multimodal and Vision-Language AI
- Python · PyTorch · Vision-Language Models · Visual Grounding · Hallucination Evaluation · Multimodal Learning
- Generative AI and LLMs
- Diffusion Models · Large Language Models (including Llama and Qwen) · Prompt Safety · Parameter-Efficient Fine-Tuning (LoRA, QLoRA)
- Computer Vision
- Object Detection · Segmentation · Pose Estimation · Image Quality Assurance
- 3D Vision
- Structure from Motion · Multi-View Stereo · COLMAP · Open3D · Neural Radiance Fields (NeRF) · 3D Gaussian Splatting
- Production ML and MLOps
- Model Serving · Deployment · Offline and Online Evaluation · A/B Testing · Dataset Design
- Backend and Cloud
- FastAPI · Docker · Nginx · Redis · Google Cloud Platform · Google Cloud Storage · SQL Server
Honors and Awards
- Champion, BASIS National ICT Awards, 2020 (Retouched.ai)
- Finalist, Asia Pacific ICT Alliance Awards, 2021
- Artificial Intelligence in Advertising, invited workshop speaker, Daffodil International University
- magna cum laude; Top 3%; Merit Scholarship and Tuition Fee Waiver, AIUB
Contact
Open to research collaboration, speaking invitations, and professional inquiries in multimodal AI, vision-language models, and applied machine learning.
- mdashikur.rafi@gmail.com
- GitHub
- Google Scholar
- Dhaka, Bangladesh