Md Ashikur Rahman is a Lead AI Engineer at The KOW Company, specializing in multimodal AI, vision-language systems, generative AI, and applied computer vision.

Research interests: trustworthy multimodal learning; vision-language models; AI safety and robustness.

He currently leads a multidisciplinary team of 15+ engineers and researchers and delivers enterprise AI products used by 200+ brands. He architected computer-vision systems that have processed more than 4.5 million images globally.

Recent Highlights

  • ICDAR 2026 oral paper on diagram understanding with LogicBench-1K (research lead and corresponding author).
  • ICCA 2026 paper on uncertainty-calibrated retrieval for medical vision-language hallucination reduction (last author).
  • Manuscripts under review on conformal risk control for LLM tool calls, detector-based grounding metrics in video QA, math-encoded jailbreaks, and residual stream rebalancing.
  • The Fitting Room: cross-brand virtual try-on platform supporting products from 170+ brands.
  • CogniX: early-stage R&D for proprietary multimodal content creation.

Education

B.Sc. in Computer Science and Engineering, American International University-Bangladesh, 2011-2015. CGPA 3.87/4.00 (WES equivalent 3.94/4.00); magna cum laude; top 3%; Merit Scholarship and Tuition Fee Waiver.

Publications

Stroke-Level Connectivity Verification: Grounding Vision-Language Models Against Topology Hallucination in Diagram Understanding · ICDAR 2026 · Oral · Research lead and corresponding author · LogicBench-1K

UCAR: Uncertainty-Calibrated Adaptive Retrieval for Hallucination Reduction in Medical Vision-Language Models · ICCA 2026 · Last author

Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls · Manuscript under review, 2026 · First author

When Detector-Based Grounding Metrics Measure Vocabulary: A Cautionary Audit of Entity Claims in Video-QA Reasoning Traces · Manuscript under review, 2026 · First author

Decoding Harm: Do Reasoning Models Resist Math-Encoded Jailbreaks? · Manuscript under review, 2026 · First author

Residual Stream Rebalancing: Training-Free Hallucination Mitigation in Vision-Language Models · Manuscript under review, 2026 · Last author

Automated Detection of Diabetic Retinopathy Using Deep Residual Learning · International Journal of Computer Applications, 2020 · First author

Full publication list on Google Scholar →

Selected Projects

Retouched.ai Object Detection and Segmentation
Lead AI Engineer Production
  • Developed salient-object segmentation for background removal; improved segmentation quality by 17% on internal benchmarks, reduced processing time by 30%, enabled uploads of up to 257 MB, and achieved a 2.27-second average processing time across standard workloads.
  • Scaled Retouched.ai to process 4.5M+ images globally for hundreds of customers using PyTorch, U²-Net-inspired salient-object segmentation, FastAPI, and Google Cloud Platform.
Omnimage.ai AI Image and Video Generation
Lead AI Engineer Production
  • Co-designed and launched image- and video-generation APIs used by 200+ brands for creative and product-image workflows.
  • Engineered workflows for reference-image conditioning, asynchronous processing, prompt classification, intent routing, and automated model selection.
The Fitting Room Cross-Brand Virtual Try-On Platform (In development)
Lead AI Engineer In development
  • Conceived and co-led a unified cross-brand virtual try-on platform supporting products from 170+ brands.
  • Architected a Dockerized FastAPI/Nginx backend using SQL Server, Google Cloud Storage, Redis, recommendation services, and 2D virtual try-on pipelines.
Enterprise AI Catalog Audit and Image Quality Assurance Platform (In development) Catalog, Image, and Content Quality Audit
Lead AI Engineer In development
  • Leading the development of an AI catalog-audit and image quality assurance platform for a major US retail client.
  • Automating image QA, catalog validation, metadata checks, and content-health monitoring across DAM and CMS workflows using NLP, PyTorch, and FastAPI.
CogniX Proprietary Multimodal Content Creation Platform (Active R&D)
Lead AI Engineer Active R&D
  • Leading early-stage R&D for multimodal product-visualization and editing workflows.
  • Researching and prototyping multi-reference fusion and editing workflows for garment replacement, scene composition, product visualization, and style transfer using PyTorch, Diffusers, ComfyUI, and parameter-efficient fine-tuning (LoRA, QLoRA).

Experience

The KOW Company Dhaka, Bangladesh
Lead AI Engineer Current Jan 2023 - Present
  • Lead a multidisciplinary team of 15+ ML engineers, software engineers, and junior researchers, driving technical strategy, system architecture, engineering execution, quality standards, and cross-functional delivery.
  • Lead the development and deployment of scalable AI systems spanning virtual try-on, computer vision, 3D reconstruction, and multimodal content creation (Retouched.ai, Omnimage.ai, HoloSnap.ai, CogniX, and The Fitting Room).
  • Direct applied research on AI hallucination, prompt safety, visual grounding, and video understanding, operationalizing research into production-ready models and evaluation frameworks.
  • Mentor and supervise four junior researchers in literature review, research planning, experimental design, reproducible implementation, result analysis, and academic writing.
  • Develop and deploy deep learning-based computer vision and image-processing systems; publish reproducible code, datasets, and model checkpoints on GitHub and Hugging Face.
Senior Machine Learning Engineer Jul 2021 - Dec 2022
  • Improved object detection and segmentation performance by 20-35% across internal evaluation benchmarks; the resulting models were later deployed in Retouched.ai.
  • Led 6+ client ML engagements from requirements gathering through production delivery; built offline evaluation pipelines and A/B testing workflows to validate model quality, inference performance, and business and operational outcomes.
Machine Learning Engineer Jul 2020 - Jun 2021
  • Built deep learning models for object recognition, image segmentation, and background-removal workflows; developed scalable preprocessing, training, and A/B testing pipelines.
Smart Technologies (BD) Ltd Dhaka, Bangladesh
Senior Software Engineer Sep 2016 - Dec 2019
  • Led .NET and SQL Server supply-chain systems on a 4 TB database, reducing report generation from 20 minutes to 40-54 seconds.
  • Achieved 70-75% process automation and 99.9% synchronization success for offline-capable enterprise workflows.
Proggasoft Dhaka, Bangladesh
Software Engineer Mar 2015 - Aug 2016
  • Developed ASP.NET MVC features, backend services, and database integrations for DevSkill.com.

Technical Skills

Multimodal and Vision-Language AI
Python · PyTorch · Vision-Language Models · Visual Grounding · Hallucination Evaluation · Multimodal Learning
Generative AI and LLMs
Diffusion Models · Large Language Models (including Llama and Qwen) · Prompt Safety · Parameter-Efficient Fine-Tuning (LoRA, QLoRA)
Computer Vision
Object Detection · Segmentation · Pose Estimation · Image Quality Assurance
3D Vision
Structure from Motion · Multi-View Stereo · COLMAP · Open3D · Neural Radiance Fields (NeRF) · 3D Gaussian Splatting
Production ML and MLOps
Model Serving · Deployment · Offline and Online Evaluation · A/B Testing · Dataset Design
Backend and Cloud
FastAPI · Docker · Nginx · Redis · Google Cloud Platform · Google Cloud Storage · SQL Server

Honors and Awards

  • Champion, BASIS National ICT Awards, 2020 (Retouched.ai)
  • Finalist, Asia Pacific ICT Alliance Awards, 2021
  • Artificial Intelligence in Advertising, invited workshop speaker, Daffodil International University
  • magna cum laude; Top 3%; Merit Scholarship and Tuition Fee Waiver, AIUB

Additional Academic Training

Individual Study in Mathematics for Research, 2025 (through The KOW Company): Multivariable Calculus with Dr. Anindita Paul; Differential Geometry with Dr. A. K. M. Nazimuddin, Department of Mathematics and Data Science, East West University.

Contact

Open to research collaboration, speaking invitations, and professional inquiries in multimodal AI, vision-language models, and applied machine learning.