Ik deploy en host je llm of ai app op de cloud

Sommige informatie is automatisch vertaald.

Sri Lanka

Ik spreek Engels

Senior Software Engineer POS, ERP en AI-gestuurde web- en mobiele oplossingen

Senior Software Engineer die schaalbare POS, ERP, MVP, web- en mobiele apps bouwt met focus op prestaties, veiligheid en onderhoudbaarheid. Bedreven in microservices, API-ontwikkeling, database-optima...
Over deze dienst

Wil je je LLM snel, geoptimaliseerd en kostenefficiënt naar productie brengen?


Je bent op de juiste plek!


Ik deploy en configureer open-source LLMs op GPU cloud servers en Kubernetes, met afgestemde inference engines voor snelle reacties en lagere GPU-kosten.


Wat ik aanbied:


  • LLM deployment (Llama, Mistral, Qwen, DeepSeek, Gemma)
  • Inference engine configuratie (vLLM, TGI, Ollama, Triton)
  • Model quantization (AWQ, GPTQ, GGUF)
  • GPU optimalisatie & kostenbesparing
  • OpenAI-compatibele LLM API setup
  • Privé & zelfgehoste LLM deployment
  • RAG & AI app backend deployment
  • Kubernetes LLM cluster met auto-scaling
  • Docker & CI/CD voor LLM apps
  • Monitoring & performance tuning


Tech stack:


Inference: vLLM || TGI || Ollama || NVIDIA Triton || TensorRT-LLM || llama.cpp || SGLang

Modellen: Hugging Face || Llama || Mistral || Qwen || DeepSeek

Cloud & GPU: AWS || Google Cloud || Azure || RunPod || Lambda Labs

Containers: Docker || Kubernetes || Helm

Monitoring: Prometheus || Grafana


Waarom voor mij kiezen?


  • Snelle levering
  • Gratis consultatie
  • Geoptimaliseerd voor snelheid & kosten
  • Volledige documentatie
  • Ondersteuning na deployment


Laten we je LLM vandaag nog live zetten!


Tools:

Kubernetes

•

Docker

•

Amazon EKS

•

Google Kubernetes Engine

Frameworks:

Npm

•

Terraform

•

Ansible

•

Pop

•

Crossplane

Cloudprovider:

Amazon Web Services

•

microsoft azure

Programmeertaal:

Bash

•

C

•

Go

•

Java

•

JavaScript

•

Lua

•

PHP

•

Python

•

Ruby

•

Golang

Expertise:

Installatie

•

Debuggen

•

Configuratie