Live evaluation, with benchmarks generated fresh on every run
A framework that generates new synthetic test data for each evaluation run, and an ablation study of when that data is a realistic stand-in for a real benchmark.
A framework that generates new synthetic test data for each evaluation run, and an ablation study of when that data is a realistic stand-in for a real benchmark.
Turning a hand-managed home server into Ansible and Docker Compose, with service routes, DNS and documentation generated from one inventory.
A rework of my bachelor-thesis air-quality application, starting with a cleaner structure for the data pipeline and a containerised database.
A Blokus game engine with a CLI and a web GUI, built by a team of four with LLM agents, and a record of which prompting guidelines held up and which failed.
A multi-agent RAG system that answers factual questions from live web evidence, and what each of its eight components actually contributes.
Classical, hybrid and graph-based recommenders compared on the X-Wines dataset, including cold-start users and wines.
A controlled comparison of nine classifiers trained on real, SMOTE-generated and deep-learning-generated obesity-risk data.
A web application and data pipeline that collects air-pollution measurements daily and tells residents how their city compares with its own history.