Directory / Braintrust
Braintrust
Verified
Evaluation platform for datasets, experiments, scoring, and logs.
Braintrust provides datasets, experiment comparison, scoring, and logging for AI products. It gives teams a shared record for changes to prompts, models, and harnesses.
- Model support
- Multi-model
- Deployment
- Cloud, Self-hosted
- Open source
- No
- License
- Proprietary (some client tooling public)
- Pricing model
- Mixed
- Added / updated
- 19 Aug 2026 / 19 Aug 2026
Assessment
Braintrust turns evaluation datasets and experiments into maintained engineering artifacts. Its experiment workflow differs from application APM, and coding-agent execution sits outside the platform. Verify self-hosting terms for the required plan.
Strengths
- First-class datasets and experiments.
- Logs associated with scores.
- Supports iterative harness evaluation.
Limitations
- Application-code writing remains outside the platform.
- Delivery outcomes require additional instrumentation.