I will build an llm evaluation framework

S
sushii_98
S
sushii_98
Sushma Sharma
Sommige informatie wordt in het Engels weergegeven.

Over deze dienst

Your AI system looks fine in demos. But is it still working 

correctly in production next week?


Most teams find out their AI broke when a user complains. 

I build evaluation infrastructure that catches failures before 

users do automatically.


What I build:

- Retrieval evaluation (are you fetching the right content?)

- Faithfulness scoring (is the answer grounded or hallucinated?)

- Regression detection (did your last prompt change break anything?)

- LLM-as-judge pipelines with no ground truth required

- Streamlit dashboard showing quality trends over time


If you are running AI in production without evals, you are 

flying blind. Message me and I will tell you where your 

system is most likely failing.

Maak kennis met Sushma Sharma

Sushma Sharma

AI Engineer

  • Afkomstig uitIndia
  • Lid sindsjun 2026
  • Gem. reactietijd3 uur
  • Talen

    Engels
AI Engineer with 5+ years of experience designing and building AI systems: 🔹 Multi-agent & agentic AI: Built ReAct-based LLM workflows for automated proposal generation. A Multi-RAG framework I designed identified $0.8M in margin leakage in customer proposals. 🔹 RAG & retrieval optimization: Improved retrieval accuracy by 20% using semantic query enhancement. Built GenAI risk assessment workflows for safety documentation. 🔹 LLMOps: Evaluation pipelines, prompt testing, and model performance monitoring for enterprise AI. Communicate well with cross-functional teams.

Mijn portfolio

Andere AI-development diensten die ik aanbied