Custom LLM Fine-Tuning & Evaluation | Case Study | Saad Ullah Bilal
← Back to Case Studies
02

Custom LLM Fine-Tuning & Evaluation

The Challenge

Off-the-shelf models were too generic and too expensive for specialized, domain-specific tasks.

What I Built

A fine-tuning and benchmarking pipeline across Llama 3.1, Mistral-7B, Phi-3.5, and Qwen using LoRA — evaluating which model delivers the best performance-to-cost ratio for each use case.

Tech Stack
TransformersLoRAHugging FacePyTorchLlama 3.1Qwen
Outcome

Matched larger-model performance at a fraction of inference cost, enabling cost-effective deployment at scale.

Governed AI Stack Layers Used