Ik bouw een production rag systeem over jouw documenten

C
chlee9
C
chlee9
Cheolhee Lee
Sommige informatie is automatisch vertaald.

Over deze dienst

Automatische vertaling

Ik bouw production RAG systemen die antwoorden geven vanuit JOUW documenten met citaties, niet met hallucinaties.


Wat je krijgt

- Documentinname pipeline (PDF, DOCX, HTML, Notion, Confluence) met layout-gevoelige chunking

- Hybride retrieval: BM25 + dense vectors + reranking, afgestemd op jouw eigen eval set

- Antwoord synthese met inline citaties en weigering wanneer bewijs ontbreekt

- FastAPI / Node service, Docker, CI en een beheerdersdashboard voor herindexering


Waarom ik

Ik lever LLM systemen in productie. Recent gemeten resultaten: 70% lagere p95 latency, 38% lagere servicekosten, 49% minder output tokens door context caching, gestructureerde output en model routing.


Stack

Python, TypeScript, LangChain, LlamaIndex, OpenAI / Anthropic / vLLM, pgvector, Qdrant, Elasticsearch, Postgres, Redis, Docker, AWS.


Stuur me een voorbeeld van je documenten en de vragen die je beantwoord wilt hebben, en ik vertel je precies wat haalbaar is voordat je bestelt.

Maak kennis met Cheolhee Lee

Cheolhee Lee

AI Full Stack Developer specializing in LLM and RAG optimization

  • Afkomstig uitZuid-Korea
  • Lid sindsapr 2021
  • Gem. reactietijd1 uur
  • Talen

    Koreaans, Engels
I keep my employer and my clients unnamed here. I ship production AI systems end to end at an undisclosed B2B AI SaaS company - a sales-automation SaaS and a public-sector AI evaluation platform. Measured: LLM p95 latency -70%, serving cost -38%, output tokens -49% via context caching and structured output. 125x list speedup, threads query 411ms to 1.6ms, bundle 21.7MB to 2.3MB. Re-homed three LLM models to an on-prem DGX with zero downtime; passed TTA review for Korea's AI Verification program. TypeScript, Python, Rust, Go, React, PostgreSQL, AWS, RAG, vLLM, MCP.

Automatische vertaling

Mijn portfolio

Andere AI-development diensten die ik aanbied