I will clean up your data
ML engineer
Over deze dienst
I clean and format raw datasets for LLM training and fine-tuning. Services include deduplication, PII redaction, noise removal, and conversion to standard formats (JSONL, Alpaca, ChatML). Python-based, NDA-compliant, and tailored to your model's architecture. Message me before ordering to discuss dataset size and requirements.
Mijn portfolio
Veelgestelde vragen
Do you sign NDAs or handle sensitive data securely?
Yes. I sign NDAs before accessing any data and delete all files upon project completion. I never use client data for my own models.
What formats do you accept and deliver?
I accept CSV, JSON, JSONL, Parquet, TXT, and SQL dumps. Delivery formats include JSONL, Parquet, Alpaca, ChatML, or custom schemas for HuggingFace/Llama.cpp.
Can you handle multilingual or code-heavy datasets?
Yes. Specify languages and coding standards upfront so I can apply appropriate tokenization, encoding fixes, and syntax validation.
How do you ensure quality after cleaning?
Every delivery includes a validation report with metrics (dedup rate, PII count, format compliance) plus 10–20 sample rows for manual review.
What if my dataset is larger than the Premium package?
Message me with your row count and requirements. I provide custom quotes for datasets exceeding 1M rows or requiring specialized processing.
Do you offer ongoing or recurring cleaning services?
Yes. Many clients need continuous pipeline support. Contact me to discuss retainer or subscription arrangements.

