Stanford, California
Stanford University
Stanford's edge is measurement and framing: it names the category (foundation models), then publishes the numbers everyone else argues over.
Key findings
- 2009
ImageNet
Fei-Fei Li and collaborators
14 million hand-labelled images across 20,000 categories, plus the annual challenge that made progress comparable.
Why it matters here · The 2012 AlexNet result on ImageNet is the moment deep learning stopped being a niche. Every image model in the catalog descends from that benchmark culture.
14M labelled images - 2015
CRISPR base and delivery engineering
Stanley Qi, Lei Stanley Qi lab and collaborators
CRISPRi/CRISPRa — using catalytically dead Cas9 to switch genes off and on without cutting DNA.
Why it matters here · The programmable-biology layer under the biotech companies now appearing in the tracker.
- 2016
SQuAD
Rajpurkar, Zhang, Lopyrev, Liang
100,000 crowd-written question–answer pairs over Wikipedia passages, with a public leaderboard.
Why it matters here · Set the benchmark-and-leaderboard culture that BERT, then GPT, were tuned against.
100k QA pairs - 2021
“Foundation models” — naming the shift
Bommasani, Liang et al., CRFM
A 200-page report arguing that one pretrained model, adapted downstream, was becoming the substrate of the whole field — and that homogenisation concentrates risk.
Why it matters here · This is the framing behind Geek Mode: nearly every product in the tracker is a thin adaptation layer over four or five base models, so one vendor's change propagates everywhere.
- 2022
HELM — holistic evaluation
Liang et al., CRFM
Evaluating models across accuracy, calibration, robustness, fairness, bias, toxicity and efficiency simultaneously, on the same scenarios.
Why it matters here · The reason we mark model claims as disclosed / inferred / unknown rather than repeating a vendor's single benchmark number.
- 2022
FlashAttention
Tri Dao, Fu, Ermon, Rudra, Ré
An IO-aware exact attention kernel that tiles computation in SRAM instead of materialising the attention matrix in HBM.
Why it matters here · Long-context pricing in the tracker only works because of this — it cut attention memory from quadratic to linear in practice.
2–4× faster training, 10–20× less memory - 2023
Alpaca — cheap instruction tuning
Taori, Gulrajani et al.
Instruction-tuned a 7B LLaMA on 52K self-generated examples for a few hundred dollars of compute.
Why it matters here · Kicked off the open fine-tune ecosystem — the reason small vendors in the catalog can ship a credible assistant without training a base model.
<$600 of compute - 2023
Direct Preference Optimization (DPO)
Rafailov, Sharma, Mitchell et al.
Showed that RLHF's reward model and PPO loop can be replaced by a single classification-style loss on preference pairs.
Why it matters here · Why small labs in the catalog can align a model at all — DPO removed the most expensive, least stable part of the alignment pipeline.
- 2023
Generative agents — the Smallville simulation
Park, O'Brien, Cai et al.
25 LLM-driven characters with memory streams, reflection and planning produced believable emergent social behaviour over simulated days.
Why it matters here · The memory/reflection loop in this paper is the template most agent frameworks we track still use.
25 agents, 2 simulated days - 2024
The inference cost collapse
HAI AI Index
Measured the cost of GPT-3.5-level performance falling by more than two orders of magnitude in under two years.
Why it matters here · Explains the pricing pattern in our change feed: entry prices fall or stay flat while capability jumps, because the underlying token cost keeps collapsing.
~280× cheaper in 18 months
On the history timeline
Milestones on the 1943 → today timeline credited to this institution.
- 1980Expert systems go commercial (XCON, MYCIN)John McDermott, Edward Shortliffe et al.Symbolic era & the winters
- 1998PageRank — ranking by link structureLarry Page & Sergey BrinConnectionist revival
- 2005Stanley wins the DARPA Grand ChallengeSebastian Thrun & the Stanford Racing TeamDeep learning boom
- 2009ImageNet — 14 million labelled imagesFei-Fei Li, Jia Deng et al.Deep learning boom
What to follow
Stanford Emerging Technology Review (SETR)
Annual review · Annual report + rolling briefs
Faculty-written state-of-the-field chapters across ten technology areas — AI, semiconductors, quantum, robotics, synthetic biology, space, energy, materials, neuroscience, cryptography — each with policy implications.
HAI AI Index Report
Index / dataset · Annual
The quantitative baseline for the field: training compute, model releases, inference cost curves, benchmark saturation, investment and policy counts.
Center for Research on Foundation Models (CRFM)
Research lab · Continuous + HELM leaderboards
HELM evaluations and the Foundation Model Transparency Index — independent checks on the model claims vendors put in their docs.
How to cite this page
Free to cite and reuse under CC BY 4.0. Permalink: https://tomorrow.aliensquad.ai/academia/stanford
Tomorrow. (2026). Stanford University — key findings and research feeds [Research tracker entry]. AlienSquad. Retrieved 2026-09-17, from https://tomorrow.aliensquad.ai/academia/stanford
@misc{tomorrow-academia-stanford,
author = {{Tomorrow}},
title = {Stanford University — key findings and research feeds},
year = {2026},
publisher = {AlienSquad},
howpublished = {\url{https://tomorrow.aliensquad.ai/academia/stanford}},
note = {Accessed: 2026-09-17}
}