Live evaluation, with benchmarks generated fresh on every run

CodeCookbookProgress reports

University of Mannheim, six-month team project. A team of three.

Problem

Models may have seen public benchmarks during training, so a good score can mean memorisation. Our framework follows a Generate, Evaluate, Trash method: it creates fresh synthetic data for every run, evaluates on it and discards it, so nothing can have been memorised. The research question is whether such data is a realistic stand-in for the real benchmark, and under which configuration.

My part

I built most of the framework and its test suite: I authored 312 of the repository’s 358 commits. That covers the pipeline, the generators, benchmark profiling, the calibration phase, the evaluators and plotting, and the taxonomy task, along with the tests (more than 1,000 test functions) and the helper scripts. The framework is about 11,000 lines of Python.

Approach

Results

Generated score against the real benchmark’s score, for the best cell of each task:

TaskBest cellGeneratedReal
Grammatical error correction (ERRANT F0.5)inverse, seeded0.330.23
Spam detection (F1)forward, seeded0.980.97
Sentiment analysis (macro-F1)inverse, seedless0.520.54
Taxonomy induction (F1)inverse, seedless0.650.82

What I’d do differently

Following the report’s AI-tool declaration, the code was written with Claude Code.

How to run it

Python and an API key for at least one model provider are needed. From the repository, install framework/requirements.txt, download the spaCy model, build the benchmark files with the scripts.benchmarks.prepare_* scripts, copy example.env to .env, and run:

python -m framework.main --config framework/configs/gec/config.yaml

The Cookbook covers every setting, the four generation cells, calibration, the judge and how to add a task.