I will optimize your llm app for lower latency and cost


Over deze dienst
I optimize LLM applications that are too slow or too expensive.
What I have done in production:
- Cut p95 latency 70 percent and serving cost 38 percent with context caching, structured output and model routing
- Reduced output tokens 49 percent without losing quality
- Sped up list endpoints 125x and a threads query from 411ms to 1.6ms
- Re-homed three LLM and embedding models to an on-premise DGX with zero downtime
How it works:
1. You share your prompts, traces, model config and your latency or cost numbers
2. I profile the pipeline and send a prioritized report
3. I implement the fixes and show before and after benchmarks
Stack: OpenAI, Anthropic, vLLM, LangChain, RAG, pgvector, Python, TypeScript, Rust, Go, AWS.
Tell me your current p95 and monthly spend and I will tell you what is realistic.
Maak kennis met Cheolhee Lee
AI Full Stack Developer specializing in LLM and RAG optimization
- Afkomstig uitZuid-Korea
- Lid sindsapr 2021
- Gem. reactietijd1 uur
Talen
Koreaans, Engels
