Ik zet je LL.M. voor productie in met vLLM op je cloud GPU

I
ike_kolawole
I
ike_kolawole
Joshua
Sommige informatie is automatisch vertaald.

Over deze dienst

Automatische vertaling

Een model dat in een notebook werkt en eentje dat 1.000 gelijktijdige gebruikers aankan, zijn twee verschillende projecten. Ik doe het tweede.


Ik ben een senior ML-infrastructuur engineer, met 7 jaar ervaring in productiebetrouwbaarheid. Recent: infrastructuur die meer dan 50.000 gelijktijdige taken afhandelt in minder dan 100 ms voor meer dan 1.000 gebruikers, gebouwd op vLLM en SGLang met GPU Kubernetes node pools, met 35% lagere infra-kosten dan toen ik begon.


Wat je krijgt:

  • Jouw open-weight model (Llama, Mistral, Qwen, DeepSeek) bediend met vLLM
  • Gehost in JOUW cloud account: jij houdt de sleutels, data en endpoint
  • GPU autoscaling en batching afgestemd op inference, niet webverkeer
  • Load-testresultaten, zodat je de echte capaciteit weet voordat je gebruikers het ontdekken
  • Monitoring dashboards; runbook op Standard/Premium

Als je GPU-budget en latency-doel niet overeenkomen, zeg ik dat voordat je een cent uitgeeft aan compute.


Basic: één model, één GPU node, load-tested.

Standard: voegt autoscaling, tuning, dashboards toe.

Premium: multi-model, kostenpass, overdracht.


Stuur je model- en trafficgegevens; het requirements formulier dekt de rest.

Maak kennis met Joshua

Joshua

Senior Platform Engineer

  • Afkomstig uitNigeria
  • Lid sindsjul 2026
  • Talen

    Engels
Deployments on my platforms are boring, and that's the point. Seven years of production infrastructure: multi-cloud Kubernetes (EKS/AKS), Terraform at 20+ module scale, CI/CD with automatic rollback at 99.9% deploy success. One team went from 40% failed deploys to 5% after I rebuilt their pipelines; a fintech's PCI-DSS audits returned zero critical findings. I also build ML serving infrastructure: 50,000+ concurrent tasks at sub-100ms using vLLM on GPU node pools. Hand me a breaking pipeline, an untrusted cluster, or a cloud bill that grew quietly, and get it fixed properly.

Automatische vertaling

Mijn portfolio