Epoch and METR launch MirrorCode to test long-horizon AI coding

MirrorCode, a benchmark co-developed with METR, asks models to reimplement open-source programs over long horizons. Early results say models can finish such work, but success depends heavily on inference budget and contamination controls.
LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: What's the largest software project AI can complete on its own?