softwarefactory.build

A directory and guide library for agent-native software factories.

Directory / Braintrust

Braintrust

Verified

Eval-first platform: datasets, experiments, and logging.

Braintrust is a platform for evaluating AI products: datasets, experiment comparison, scoring, and logging. Teams use it as the system of record for “did this harness get better.”

Model independence
Multi-model
Deployment
Cloud, Self-hosted
Open source
No
License
Proprietary (some client tooling public)
Pricing model
Mixed
Added / updated
19 Aug 2026 / 19 Aug 2026

www.braintrust.dev · Docs

Editorial take

Braintrust is what you buy when eval is the product discipline. It is closer to a lab notebook than to Datadog. Self-host is an enterprise question — verify. Not an agent.

Strengths

  • Experiments and datasets as first-class objects.
  • Logging tied to scores, not only traces.
  • Used as a harness-improvement loop.

Limitations

  • Commercial platform.
  • Will not write application code.
  • Easy to build a beautiful eval that ignores revert rate.

Fields last checked against primary sources on . Pricing and plan names rot; follow the vendor URL.