Off-the-shelf models were too generic and too expensive for specialized, domain-specific tasks.
A fine-tuning and benchmarking pipeline across Llama 3.1, Mistral-7B, Phi-3.5, and Qwen using LoRA — evaluating which model delivers the best performance-to-cost ratio for each use case.
Matched larger-model performance at a fraction of inference cost, enabling cost-effective deployment at scale.