Filtered lists / Verification gates
Verification gates
Review, eval, and merge tools that can gate a change. Autocomplete products are excluded unless they document that role.
Inclusion follows the query’s structured fields; this is not a score. 22 matching listings.
- Braintrust
Evaluation platform for datasets, experiments, scoring, and logs.
- CodeQL
GitHub semantic code analysis with query-based CI checks.
- CodeRabbit
AI pull-request review for GitHub and GitLab.
- Cubic
AI code review for GitHub pull requests.
- Dagger
Programmable containerized CI/CD engine for agent workflows.
- DeepEval
Pytest-oriented evaluation library for LLM systems.
- DeepSource
Hosted static analysis with Autofix pull requests.
- Galileo
Enterprise evaluation and observability for AI systems.
- Giskard
Open testing and red-teaming platform for AI applications.
- Graphite
Stacked pull requests with CLI workflow and AI review.
- Greptile
AI pull-request review with repository-wide context.
- Inspect
UK AISI open framework for evaluating language-model agents.
- Langfuse
Self-hostable open-source LLM and agent observability.
- LangSmith
LangChain commercial tracing, evaluation, and prompt platform.
- LangWatch
LLM observability and evaluation platform with open components.
- Mastra
TypeScript agent framework with workflows and evaluation hooks.
- Mergify
GitHub merge queue and pull-request automation rules.
- Phoenix
Arize open-source observability and evaluation for LLM applications.
- Promptfoo
Open evaluation and red-team framework for LLM applications.
- Ragas
Open evaluation metrics for RAG and LLM applications.
- Semgrep
Teachably fast static analysis with optional AI-assisted rules.
- Trunk
Unified linters, formatters, and CI reliability checks.