Academia

Stanford, California

Stanford University

Stanford's edge is measurement and framing: it names the category (foundation models), then publishes the numbers everyone else argues over.

AIPolicyQuantumBiotechOfficial site
5 key findings4 timeline milestones3 sources to follow

Key findings

  1. 2021

    “Foundation models” — naming the shift

    Bommasani, Liang et al., CRFM

    A 200-page report arguing that one pretrained model, adapted downstream, was becoming the substrate of the whole field — and that homogenisation concentrates risk.

    Why it matters here · This is the framing behind Geek Mode: nearly every product in the tracker is a thin adaptation layer over four or five base models, so one vendor's change propagates everywhere.

  2. 2009

    ImageNet

    Fei-Fei Li and collaborators

    14 million hand-labelled images across 20,000 categories, plus the annual challenge that made progress comparable.

    Why it matters here · The 2012 AlexNet result on ImageNet is the moment deep learning stopped being a niche. Every image model in the catalog descends from that benchmark culture.

    14M labelled images
  3. 2022

    HELM — holistic evaluation

    Liang et al., CRFM

    Evaluating models across accuracy, calibration, robustness, fairness, bias, toxicity and efficiency simultaneously, on the same scenarios.

    Why it matters here · The reason we mark model claims as disclosed / inferred / unknown rather than repeating a vendor's single benchmark number.

  4. 2024

    The inference cost collapse

    HAI AI Index

    Measured the cost of GPT-3.5-level performance falling by more than two orders of magnitude in under two years.

    Why it matters here · Explains the pricing pattern in our change feed: entry prices fall or stay flat while capability jumps, because the underlying token cost keeps collapsing.

    ~280× cheaper in 18 months
  5. 2023

    Alpaca — cheap instruction tuning

    Taori, Gulrajani et al.

    Instruction-tuned a 7B LLaMA on 52K self-generated examples for a few hundred dollars of compute.

    Why it matters here · Kicked off the open fine-tune ecosystem — the reason small vendors in the catalog can ship a credible assistant without training a base model.

    <$600 of compute

On the history timeline

Milestones on the 1943 → today timeline credited to this institution.

  • 1980Expert systems go commercial (XCON, MYCIN)John McDermott, Edward Shortliffe et al.Symbolic era & the winters
  • 1998PageRank — ranking by link structureLarry Page & Sergey BrinConnectionist revival
  • 2005Stanley wins the DARPA Grand ChallengeSebastian Thrun & the Stanford Racing TeamDeep learning boom
  • 2009ImageNet — 14 million labelled imagesFei-Fei Li, Jia Deng et al.Deep learning boom

What to follow