Research
W2–W10 closure: EvaluationResultV2 (10-dim + 6-level), Benchmark v2 (100 tasks), Ablation A–F, Significance (bootstrap CI / McNemar / Wilcoxon), Reproducibility L0–L5, Failure F01–F15. See docs/benchmark.md and research/paper/V2_paper_draft.md.
Recent experiments
No results yet — run
uv run python research/experiments/run_ablation.py --limit 20 --out research/resultsRunner provenance: ablation configs at
research/experiments/ablation_matrix.py (A LLM-only → F Full). Significance: packages/evaluation/src/dsa_evaluation/significance.py.