Md Ashikur Rahman is a Lead AI Engineer at The KOW Company, specializing in multimodal AI, vision-language systems, generative AI, and applied computer vision.

Research interests: trustworthy multimodal machine learning; vision-language models; AI safety and robustness.

He currently leads a multidisciplinary team of 15+ engineers and researchers and delivers enterprise AI products used by 200+ brands. He architected computer-vision systems that have processed more than 4.5 million images globally.

Recent Highlights

  • ICDAR 2026 oral paper on diagram understanding with LogicBench-1K (last and corresponding author).
  • ICCA 2026 paper on uncertainty-calibrated retrieval for medical vision-language hallucination reduction (last author).
  • Manuscripts under review on conformal risk control for LLM tool calls, detector-based grounding metrics in video QA, math-encoded jailbreaks, and residual stream rebalancing.
  • The Fitting Room: cross-brand virtual try-on platform supporting products from 170+ brands.
  • CogniX: early-stage R&D for proprietary multimodal content creation.

Education

B.Sc. in Computer Science and Engineering, American International University-Bangladesh, 2011-2015. CGPA 3.87/4.00 (WES equivalent 3.94/4.00); magna cum laude; top 3%; Merit Scholarship and Tuition Fee Waiver.

Publications

Stroke-Level Connectivity Verification: Grounding Vision-Language Models Against Topology Hallucination in Diagram Understanding · ICDAR 2026 · Oral · Last and corresponding author · LogicBench-1K

UCAR: Uncertainty-Calibrated Adaptive Retrieval for Hallucination Reduction in Medical Vision-Language Models · ICCA 2026 · Last author

Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls · Manuscript under review, 2026 · First author

When Detector-Based Grounding Metrics Measure Vocabulary: A Cautionary Audit of Entity Claims in Video-QA Reasoning Traces · Manuscript under review, 2026 · First author

Decoding Harm: Do Reasoning Models Resist Math-Encoded Jailbreaks? · Manuscript under review, 2026 · First author

Residual Stream Rebalancing: Training-Free Hallucination Mitigation in Vision-Language Models · Manuscript under review, 2026 · Last author

Automated Detection of Diabetic Retinopathy Using Deep Residual Learning · International Journal of Computer Applications, 2020 · First author

Full publication list on Google Scholar →

Selected Projects

Retouched.ai Object Detection and Segmentation
Lead AI Engineer Production
  • Developed salient-object segmentation for background removal; improved segmentation quality by 17% on internal benchmarks, reduced processing time by 30%, enabled uploads of up to 257 MB, and achieved a 2.27-second average processing time across standard workloads.
  • Scaled Retouched.ai to process 4.5M+ images globally for hundreds of customers using PyTorch, U²-Net-inspired salient-object segmentation, FastAPI, and Google Cloud Platform.
Omnimage.ai AI Image and Video Generation
Lead AI Engineer Production
  • Co-designed and launched image- and video-generation APIs used by 200+ brands for creative and product-image workflows.
  • Engineered workflows for reference-image conditioning, asynchronous processing, prompt classification, intent routing, and automated model selection.
The Fitting Room Cross-Brand Virtual Try-On Platform (In development)
Lead AI Engineer In development
  • Conceived and co-led a unified cross-brand virtual try-on platform supporting products from 170+ brands.
  • Architected a Dockerized FastAPI/Nginx backend using SQL Server, Google Cloud Storage, Redis, recommendation services, and 2D virtual try-on pipelines.
Enterprise AI Catalog Audit and Image Quality Assurance Platform (In development) Catalog, Image, and Content Quality Audit
Lead AI Engineer In development
  • Leading the development of an AI catalog-audit and image quality assurance platform for a major US retail client.
  • Automating image QA, catalog validation, metadata checks, and content-health monitoring across DAM and CMS workflows using NLP, PyTorch, and FastAPI.
CogniX Proprietary Multimodal Content Creation Platform (Active R&D)
Lead AI Engineer Active R&D
  • Leading early-stage R&D for multimodal product-visualization and editing workflows.
  • Researching and prototyping multi-reference fusion and editing workflows for garment replacement, scene composition, product visualization, and style transfer using PyTorch, Diffusers, ComfyUI, and parameter-efficient fine-tuning (LoRA, QLoRA).

Experience

The KOW Company Dhaka, Bangladesh
Lead AI Engineer Current Jan 2023 - Present
  • Lead a multidisciplinary team of 15+ ML engineers, software engineers, and junior researchers, driving technical strategy, system architecture, engineering execution, quality standards, and cross-functional delivery.
  • Lead the design, development, and deployment of scalable AI systems spanning virtual try-on, computer vision, 3D reconstruction, and multimodal content creation (Retouched.ai, Omnimage.ai, HoloSnap.ai, CogniX, and The Fitting Room).
  • Direct applied research on AI hallucination, prompt safety, visual grounding, and video understanding, operationalizing research into production-ready models and evaluation frameworks.
  • Mentor and supervise four researchers on literature review, research planning, experimental design, reproducible implementation, result analysis, and academic writing.
  • Develop and deploy deep learning–based computer vision and image-processing systems; publish reproducible code, datasets, and model checkpoints on GitHub and Hugging Face.
Senior Machine Learning Engineer Jul 2021 - Dec 2022
  • Improved object detection and segmentation performance by 20-35% across internal evaluation benchmarks; the resulting models were later deployed in Retouched.ai.
  • Led 6+ client ML engagements from requirements gathering through production delivery; built offline evaluation pipelines and A/B testing workflows to validate model quality, inference performance, and business and operational outcomes.
Machine Learning Engineer Jul 2020 - Jun 2021
  • Built deep learning models for object recognition, image segmentation, and background-removal workflows; developed scalable preprocessing, training, and A/B testing pipelines.
Smart Technologies (BD) Ltd Dhaka, Bangladesh
Senior Software Engineer Sep 2016 - Dec 2019
  • Led .NET and SQL Server supply-chain systems on a 4 TB database, reducing report generation from 20 minutes to 40-54 seconds.
  • Achieved 70-75% process automation and 99.9% synchronization success for offline-capable enterprise workflows.
Proggasoft Dhaka, Bangladesh
Software Engineer Mar 2015 - Aug 2016
  • Developed ASP.NET MVC features, backend services, and database integrations for DevSkill.com.

Technical Skills

Multimodal and Vision-Language AI
Python · PyTorch · Vision-Language Models · Visual Grounding · Hallucination Evaluation · Multimodal Learning
Generative AI and LLMs
Diffusion Models · Large Language Models (including Llama and Qwen) · Prompt Safety · Parameter-Efficient Fine-Tuning (LoRA, QLoRA)
Computer Vision
Object Detection · Segmentation · Pose Estimation · Image Quality Assurance
3D Vision
Structure from Motion · Multi-View Stereo · COLMAP · Open3D · Neural Radiance Fields (NeRF) · 3D Gaussian Splatting
Production ML and MLOps
Model Serving · Deployment · Offline and Online Evaluation · A/B Testing · Dataset Design
Backend and Cloud
FastAPI · Docker · Nginx · Redis · Google Cloud Platform · Google Cloud Storage · SQL Server

Honors and Awards

  • Champion, BASIS National ICT Awards, 2020 (Retouched.ai)
  • Finalist, Asia Pacific ICT Alliance Awards, 2021
  • Artificial Intelligence in Advertising, invited workshop speaker, Daffodil International University
  • magna cum laude; Top 3%; Merit Scholarship and Tuition Fee Waiver, AIUB

Contact

Open to research collaboration, speaking invitations, and professional inquiries in multimodal AI, vision-language models, and applied machine learning.