Live evaluation, with benchmarks generated fresh on every run
A framework that generates new synthetic test data for each evaluation run, and an ablation study of when that data is a realistic stand-in for a real benchmark.
Machine learning engineer
I am finishing an MSc in Data Science at the University of Mannheim, with primary interests in LLM's, Evaluation and Benchmarking, DevOps and Self-Hosting.
Looking for a full-time job, remote or located in Prague, ready to start from January 2027.
A framework that generates new synthetic test data for each evaluation run, and an ablation study of when that data is a realistic stand-in for a real benchmark.
A Blokus game engine with a CLI and a web GUI, built by a team of four with LLM agents, and a record of which prompting guidelines held up and which failed.
A multi-agent RAG system that answers factual questions from live web evidence, and what each of its eight components actually contributes.