Academia

Princeton, New Jersey

Princeton University

Princeton is the field's fact-checker: rigorous work on what AI systems can and cannot do, and on the evaluations that mislead.

3 key findings1 timeline milestones2 sources to follow

Key findings

  1. 2023

    SWE-bench

    Carlos Jimenez, Ofir Press, Karthik Narasimhan et al.

    2,294 real GitHub issues from twelve Python repos; a model must produce a patch that passes the project's own tests.

    Why it matters here · The benchmark every coding agent in this catalog now quotes. When a vendor claims an agent 'resolves issues autonomously', this is the number to ask for.

    2,294 real-world issues
  2. 2024

    AI Snake Oil

    Arvind Narayanan and Sayash Kapoor

    Separated genuinely capable generative systems from predictive-AI products that cannot work, and documented widespread evaluation leakage.

    Why it matters here · Why claims in this tracker carry a confidence level and a source rather than being restated as fact.

  3. 2024

    AI agents that matter

    Kapoor, Stroebl, Narayanan et al.

    Showed agent leaderboards ignore cost, so trivially expensive baselines can top them; proposed joint accuracy-cost evaluation.

    Why it matters here · The reason we pair capability with entry price everywhere in the catalog instead of ranking on capability alone.

On the history timeline

Milestones on the 1943 → today timeline credited to this institution.

  • 2009ImageNet — 14 million labelled imagesFei-Fei Li, Jia Deng et al.Deep learning boom

What to follow