Every level

Topics at every level

Five topics cut across all the levels. Each has its own set of pages.

Every level

Who decides the next stepThese topics apply whichever level you use.
Topics

Five topics that apply at every level

Evals

Sourced

Measuring whether a change made the results better.

Evaluation frameworks

Sourced

The tools that run test sets and graders for you, and what to check before trusting their numbers.

Changing the model

Sourced

Fine-tuning, distillation, synthetic data and automated prompt tuning.

Fine-tuning and adapters

Sourced

Training a model further on your own examples, in full or with small adapters such as LoRA.

Distillation

Sourced

Training a smaller model to reproduce what a larger one does on your task.

Synthetic data

Sourced

Using a model to write training or test examples, and checking them before they are used.

Prompt optimization

Sourced

Letting a program search for better prompts against a test set.

Safety, privacy and governance

Sourced

Prompt injection, permissions, data handling and audit.

Guardrails

Sourced

Checks on what goes into a model and what comes out, and the limits of those checks.

Red teaming

Sourced

Attacking your own system on purpose, before someone else does, and turning what you find into tests.

Operations

Sourced

Cost, speed, monitoring and running models on your own hardware.

Observability

Sourced

Recording what each run did, so a bad result can be traced to the step that caused it.

AI gateways

Sourced

One entry point in front of several model providers, for keys, routing, limits, fallback and logs.

Cost optimization

Sourced

Spending fewer tokens and less time for the same result: caching, batching, smaller models, shorter context.

Running models locally

Sourced

Running open-weight models on your own hardware: what fits, quantization, and what you give up.

Working with a model

Sourced

How to brief a model, review its work and decide what to hand over.

Briefing: saying what you want

Sourced

Saying what you want clearly enough that the model does not have to guess.

Reviewing work you did not do

Sourced

Checking work you did not do yourself before it goes anywhere.

Deciding what to hand over

Sourced

Deciding which parts of a task to hand to a model and which to keep.

Calibrating trust

Sourced

Learning, from results over time, how much to rely on a model without checking.

Out there

Named products, tools and models

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Tools for building it34
  • AI Guardrails (Lakera Guard)Check Point · prompt-injection filter · formerly Lakera Guard
  • Axolotlopen source · fine-tuning library
  • BraintrustBraintrust · eval platform
  • DeepEvalConfident AI · eval framework
  • distilabelArgilla · synthetic data
  • DSPyStanford NLP · prompt programs and optimizers
  • garakNVIDIA · LLM vulnerability scanner
  • Guardrails AIGuardrails AI · guardrails framework
  • HeliconeHelicone · tracing and cost tracking
  • InspectUK AI Security Institute · eval framework
  • LangfuseClickHouse · tracing and cost tracking
  • LangSmithLangChain · eval and tracing platform
  • LiteLLMBerriAI · one API for many models
  • Llama Guard 4Meta · safety classifier · formerly Llama Guard
  • llama.cppopen source · runs models locally
  • LM StudioElement Labs · runs models locally
  • lm-evaluation-harnessEleutherAI · benchmark runner
  • MLXApple · training and inference on Apple hardware
  • NeMo GuardrailsNVIDIA · guardrails framework
  • OllamaOllama · runs models locally
  • OpenAI EvalsOpenAI · eval frameworkRetires 2026-11-30
  • OpenAI fine-tuningOpenAI · hosted fine-tuningRetired 2026
  • OpenRouterOpenRouter · one API for many models
  • OpenTelemetryopen standard · tracing standard
  • PEFTHugging Face · LoRA and other adapters
  • PhoenixArize AI · AI observability and evaluation
  • promptfoopromptfoo · eval runner
  • Ragasopen source · evals for retrieval
  • SGLangSGLang community · model serving runtime
  • Together AI fine-tuningTogether AI · hosted fine-tuning
  • TransformersHugging Face · model library
  • TRLHugging Face · fine-tuning library
  • UnslothUnsloth · fine-tuning library
  • vLLMopen source · model server

Pages at this level last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page