Door categorieën bladeren
Ontdekken
Fiverr Pro
Nederlands
$
USD
I will build an llm evaluation framework


Sushma Sharma
Over deze dienst
Your AI system looks fine in demos. But is it still working
correctly in production next week?
Most teams find out their AI broke when a user complains.
I build evaluation infrastructure that catches failures before
users do automatically.
What I build:
- Retrieval evaluation (are you fetching the right content?)
- Faithfulness scoring (is the answer grounded or hallucinated?)
- Regression detection (did your last prompt change break anything?)
- LLM-as-judge pipelines with no ground truth required
- Streamlit dashboard showing quality trends over time
If you are running AI in production without evals, you are
flying blind. Message me and I will tell you where your
system is most likely failing.
Maak kennis met Sushma Sharma
Sushma Sharma
AI Engineer
- Afkomstig uitIndia
- Lid sindsjun 2026
- Gem. reactietijd3 uur
Talen
Engels
AI Engineer with 5+ years of experience designing and building AI systems:
🔹 Multi-agent & agentic AI: Built ReAct-based LLM workflows for automated proposal generation. A Multi-RAG framework I designed identified $0.8M in margin leakage in customer proposals.
🔹 RAG & retrieval optimization: Improved retrieval accuracy by 20% using semantic query enhancement. Built GenAI risk assessment workflows for safety documentation.
🔹 LLMOps: Evaluation pipelines, prompt testing, and model performance monitoring for enterprise AI.
Communicate well with cross-functional teams.
