Layer / What must pass
Verification
Tests, reviews, hooks, and merge gates that determine whether a change is ready.
- Amazon Q Developer
AWS coding assistant with IDE, CLI, and agent features.
- Braintrust
Evaluation platform for datasets, experiments, scoring, and logs.
- Claude Code
Anthropic coding agent for terminal, IDE, desktop, and browser workflows.
- CodeQL
GitHub semantic code analysis with query-based CI checks.
- CodeRabbit
AI pull-request review for GitHub and GitLab.
- Cubic
AI code review for GitHub pull requests.
- Cursor
Agent-focused IDE with cloud agents running in isolated VMs.
- Dagger
Programmable containerized CI/CD engine for agent workflows.
- DeepEval
Pytest-oriented evaluation library for LLM systems.
- DeepSource
Hosted static analysis with Autofix pull requests.
- Devin
Cognition cloud coding agent with Devin Desktop and CLI surfaces.
- Ellipsis
AI pull-request reviewer that can open follow-up fixes.
- Factory
Commercial agent platform for software delivery workflows.
- Galileo
Enterprise evaluation and observability for AI systems.
- Giskard
Open testing and red-teaming platform for AI applications.
- GitHub Copilot
GitHub coding assistant with an issue-to-pull-request agent.
- Graphite
Stacked pull requests with CLI workflow and AI review.
- Greptile
AI pull-request review with repository-wide context.
- Inspect
UK AISI open framework for evaluating language-model agents.
- Langfuse
Self-hostable open-source LLM and agent observability.
- LangSmith
LangChain commercial tracing, evaluation, and prompt platform.
- LangWatch
LLM observability and evaluation platform with open components.
- Mastra
TypeScript agent framework with workflows and evaluation hooks.
- Mergify
GitHub merge queue and pull-request automation rules.
- Phoenix
Arize open-source observability and evaluation for LLM applications.
- Promptfoo
Open evaluation and red-team framework for LLM applications.
- Qodo
Code generation and review agents with test-aware workflows.
- Ragas
Open evaluation metrics for RAG and LLM applications.
- Semgrep
Teachably fast static analysis with optional AI-assisted rules.
- Sourcery
Refactoring and review assistant with IDE and CI surfaces.
- SWE-agent
Princeton open-source agent for SWE-bench-style issue patches.
- Trevize
Shared cloud workspaces for agent runs, previews, and pull requests.
- Trunk
Unified linters, formatters, and CI reliability checks.