Princeton, New Jersey
Princeton University
Princeton is the field's fact-checker: rigorous work on what AI systems can and cannot do, and on the evaluations that mislead.
Key findings
- 2023
SWE-bench
Carlos Jimenez, Ofir Press, Karthik Narasimhan et al.
2,294 real GitHub issues from twelve Python repos; a model must produce a patch that passes the project's own tests.
Why it matters here · The benchmark every coding agent in this catalog now quotes. When a vendor claims an agent 'resolves issues autonomously', this is the number to ask for.
2,294 real-world issues - 2024
AI Snake Oil
Arvind Narayanan and Sayash Kapoor
Separated genuinely capable generative systems from predictive-AI products that cannot work, and documented widespread evaluation leakage.
Why it matters here · Why claims in this tracker carry a confidence level and a source rather than being restated as fact.
- 2024
AI agents that matter
Kapoor, Stroebl, Narayanan et al.
Showed agent leaderboards ignore cost, so trivially expensive baselines can top them; proposed joint accuracy-cost evaluation.
Why it matters here · The reason we pair capability with entry price everywhere in the catalog instead of ranking on capability alone.
On the history timeline
Milestones on the 1943 → today timeline credited to this institution.
- 2009ImageNet — 14 million labelled imagesFei-Fei Li, Jia Deng et al.Deep learning boom