← Segments

Infrastructure & Serving

Model registries, inference platforms, vector stores and the eval/observability layer everything else is built on.

17 tools · 0 changes / 30d

Baseten

Baseten

Production inference platform with dedicated deployments.

Version
2025
Cost
Usage-based, enterprise tiers
Model
open weights
Braintrust

Braintrust

Eval-first platform for shipping LLM features safely.

Version
2025
Cost
Free tier, usage-based
Model
model-agnostic
Chroma

Chroma

Embedded-first vector store popular in prototypes and agents.

Version
1.x
Cost
Open source; cloud usage-based
Model
embeddings-agnostic
Fireworks AI

Fireworks AI

Low-latency serving for open models and custom fine-tunes.

Version
2025
Cost
Per-token usage
Model
Llama
Hugging Face

Hugging Face

The model and dataset registry the open ecosystem runs on.

Version
Hub
Cost
Free, $9/mo Pro, enterprise from $20/user
Model
hosts most open weights
LangSmith

LangChain

Tracing, evals and monitoring for LLM applications.

Version
2025
Cost
Free dev tier, $39/user/mo Plus
Model
model-agnostic
LM Studio

LM Studio

Desktop app for running and serving local models.

Version
0.3x
Cost
Free for personal use
Model
open weights
Modal

Modal Labs

Serverless GPU compute for Python-native AI workloads.

Version
2025
Cost
Per-second GPU billing, $30 free
Model
bring your own
Ollama

Ollama

Local model runner that made desktop inference trivial.

Version
0.x
Cost
Free
Model
Llama
OpenRouter

OpenRouter

One API in front of 400+ models with live price and latency data.

Version
2025
Cost
Pass-through pricing + 5%
Model
400+ models
Pinecone

Pinecone

Managed vector database built for retrieval at scale.

Version
2025
Cost
Free tier, usage-based serverless
Model
embeddings-agnostic
Replicate

Replicate

Run and ship any open model behind a simple API.

Version
2025
Cost
Per-second GPU billing
Model
FLUX
Scale AI

Scale AI

Data engine and evaluation for frontier model builders.

Version
2025
Cost
Enterprise quoted
Model
serves frontier labs
Together AI

Together AI

Open-model inference and fine-tuning at frontier speed.

Version
2025
Cost
Per-token, from $0.06/M
Model
Llama
vLLM

vLLM (PyTorch Foundation)

The de facto open-source inference engine for LLM serving.

Version
0.1x
Cost
Open source
Model
serves any open weights
Weaviate

Weaviate

Open-source vector database with built-in hybrid search.

Version
1.3x
Cost
Open source; cloud from $25/mo
Model
embeddings-agnostic

Experiment tracking and Weave evals for AI teams.

Version
2025
Cost
Free personal, from $50/user/mo
Model
framework-agnostic

Recent changes in Infrastructure & Serving

No changes recorded yet.