- AI Engineer working with deep learning, speech systems, and large language models.
- Currently pursuing Data Science at Air University Islamabad, in my 7th semester.
- Started working professionally in AI during my 5th semester.
- I work across model architecture, training, fine-tuning, optimization, inference, and deployment.
- Strong focus on Speech-to-Text, Text-to-Speech, speech translation, and low-resource languages.
- Worked extensively with Pakistani Pashto and challenging real-world speech data.
- Experienced with LLM training, fine-tuning, embeddings, synthetic data generation, and model adaptation.
- Build production AI systems using FastAPI, Docker, Kafka, CUDA, and GPU inference pipelines.
- Interested in model optimization, inference engines, Apple Silicon, GPU computing, and open-source AI infrastructure.
- Also involved in mentoring junior engineers and interns in machine learning and deep learning.
|
Denoising Telerik RadCaptcha: A Comparative Evaluation of the Effectiveness of Pre-Processing Techniques and Deep Learning Methods Using a Novel Dataset Talha Bin Omar, Tahir Sher, Abdul Rehman, M. Haroon Khan ICCK Transactions on Advanced Computing and Systems • 2026 Developed and evaluated a deep learning based approach for breaking Telerik RadCaptcha using a novel dataset of 3,000 real-world CAPTCHA images. 97.60% character accuracy • 92.08% full CAPTCHA accuracy |
- Automatic Speech Recognition
- Text-to-Speech
- Speech-to-Speech Translation
- Low-resource language modeling
- Whisper fine-tuning and optimization
- Voice and acoustic modeling
- Neural audio pipelines
- Streaming and real-time inference
- LLM fine-tuning
- Parameter-efficient training
- LoRA and DoRa
- Synthetic data generation
- Model evaluation
- Quantization
- Local LLM deployment
- Agentic AI systems
- GPU inference optimization
- CUDA workloads
- Quantized inference
- CTranslate2
- Apple Silicon / MPS
- Memory optimization
- Latency optimization
- Production inference pipelines
- FastAPI
- Docker
- Kafka
- REST APIs
- Model serving
- GPU deployments
- Data pipelines
- Git and GitHub
Built custom speech pipelines involving technologies such as:
Whisper → Neural Adapters → G2P / Linguistic Processing → TTS
Worked on speech recognition and translation under heavily degraded GSM and telephony audio, where noise, compression, low bandwidth, and unclear speech make conventional models struggle.
My work includes:
- Custom Whisper fine-tuning
- Encoder adaptation using LoRA
- Decoder hidden-state extraction
- Neural representation adapters
- Transformer and CTC architectures
- Temporal convolution bridges
- NeMo-based linguistic processing
- VITS-based speech synthesis
- End-to-end GPU inference pipelines
Developed an English-to-Pashto speech-to-speech research pipeline combining:
Whisper + Neural Adapters + NeMo G2P + Pashto VITS
Evaluation on an RTX 3090:
Generated Audio: 11.006 seconds
End-to-End Inference: 4.436 seconds
Real-Time Factor: 0.403
The complete pipeline operates significantly faster than real time.
Working on an experimental Apple Silicon MPS backend contribution for OpenNMT/CTranslate2.
The work involves extending a high-performance inference library beyond its traditional CPU and CUDA execution paths and evaluating:
- MPS operator support
- Model correctness
- WER / BLEU preservation
- Performance
- Memory behavior
- Apple Silicon compatibility
This work involves lower-level model execution and inference infrastructure rather than treating ML frameworks as black boxes.
Built Vibify, an AI-powered music recommendation platform using:
FastAPI • OpenAI Embeddings • Pinecone • Python
Implemented semantic music recommendation using vector embeddings and similarity search.
interests = [
"Speech AI",
"Large Language Models",
"Low-Resource Languages",
"Real-Time AI",
"Model Optimization",
"Quantization",
"Inference Engines",
"Apple Silicon AI",
"GPU Computing",
"Open Source AI"
]I particularly enjoy understanding how models work internally, modifying architectures, optimizing inference paths, and turning research models into production systems.
Building AI systems from model architecture to production inference.
