Best of / Verification gates
Verification gates
Review, eval, and merge tools that can sit on the factory exit. Autocomplete products that also “run tests” are not enough to appear here.
Curated by filter, not by score. A listing appears here only if its structured fields match the query. 22 listings.
- Braintrust
Eval-first platform: datasets, experiments, and logging.
- CodeQL
GitHub’s semantic code analysis, query-based and CI-native.
- CodeRabbit
High-volume AI pull-request review as a product.
- Cubic
AI code review product aimed at merging faster with fewer comments.
- Dagger
Programmable CI/CD engine agents can call as a tool.
- DeepEval
Pytest-flavored evals for LLM systems.
- DeepSource
Static analysis SaaS with Autofix, now talking to agents.
- Galileo
Evaluation and observability for production AI systems.
- Giskard
Open testing and red-teaming for AI applications.
- Graphite
Stacked PRs plus AI review, aimed at a faster merge path.
- Greptile
AI review that claims to read the whole repo, not the diff alone.
- Inspect
UK AISI’s open framework for evaluating language model agents.
- Langfuse
Open-source LLM observability you can self-host.
- LangSmith
LangChain’s commercial tracing, eval, and prompt hub.
- LangWatch
Observability and evaluation platform for LLM products.
- Mastra
TypeScript agent framework with workflows and evals attached.
- Mergify
Merge queue and automation rules for GitHub.
- Phoenix
Arize’s open observability and eval notebook for LLM apps.
- Promptfoo
Open evals and red-team for LLM apps and agents.
- Ragas
Open metrics for RAG and LLM application quality.
- Semgrep
Fast static analysis you can teach, including AI-assisted rules.
- Trunk
Unified linters, formatters, and CI flakiness as a merge gate.