Layer / What must pass
Verification
Tests, reviews, hooks, and merge gates that decide whether output is software or just a plausible patch.
- Amazon Q Developer
AWS’s coding assistant and agent, billed into the AWS account.
- Braintrust
Eval-first platform: datasets, experiments, and logging.
- Claude Code
Anthropic's coding agent across terminal, IDE, desktop, and browser.
- CodeQL
GitHub’s semantic code analysis, query-based and CI-native.
- CodeRabbit
High-volume AI pull-request review as a product.
- Cubic
AI code review product aimed at merging faster with fewer comments.
- Cursor
An agent-native IDE with cloud agents that run in isolated VMs.
- Dagger
Programmable CI/CD engine agents can call as a tool.
- DeepEval
Pytest-flavored evals for LLM systems.
- DeepSource
Static analysis SaaS with Autofix, now talking to agents.
- Devin
Cognition’s cloud software engineer; Windsurf is now Devin Desktop.
- Ellipsis
AI reviewer and fixer that can open follow-up PRs.
- Factory
Droids across the SDLC, sold as a software factory.
- Galileo
Evaluation and observability for production AI systems.
- Giskard
Open testing and red-teaming for AI applications.
- GitHub Copilot
Incumbent coding assistant with a cloud coding agent on GitHub.
- Graphite
Stacked PRs plus AI review, aimed at a faster merge path.
- Greptile
AI review that claims to read the whole repo, not the diff alone.
- Inspect
UK AISI’s open framework for evaluating language model agents.
- Langfuse
Open-source LLM observability you can self-host.
- LangSmith
LangChain’s commercial tracing, eval, and prompt hub.
- LangWatch
Observability and evaluation platform for LLM products.
- Mastra
TypeScript agent framework with workflows and evals attached.
- Mergify
Merge queue and automation rules for GitHub.
- Phoenix
Arize’s open observability and eval notebook for LLM apps.
- Promptfoo
Open evals and red-team for LLM apps and agents.
- Qodo
Generation and review agents aimed at test-aware pull requests.
- Ragas
Open metrics for RAG and LLM application quality.
- Semgrep
Fast static analysis you can teach, including AI-assisted rules.
- Sourcery
Refactoring and review assistant with a long IDE history.
- SWE-agent
Princeton’s agent that made SWE-bench a product category.
- Trevize
Shared cloud workspaces where teams and agents close the loop.
- Trunk
Unified linters, formatters, and CI flakiness as a merge gate.