softwarefactory.build

A directory and guide library for agent-native software factories.

Best of / Verification gates

Verification gates

Review, eval, and merge tools that can sit on the factory exit. Autocomplete products that also “run tests” are not enough to appear here.

Curated by filter, not by score. A listing appears here only if its structured fields match the query. 22 listings.

  1. Braintrust

    Eval-first platform: datasets, experiments, and logging.

    Verified

  2. CodeQL

    GitHub’s semantic code analysis, query-based and CI-native.

    Verified

  3. CodeRabbit

    High-volume AI pull-request review as a product.

    Verified

  4. Cubic

    AI code review product aimed at merging faster with fewer comments.

    Verified

  5. Dagger

    Programmable CI/CD engine agents can call as a tool.

    Verified

  6. DeepEval

    Pytest-flavored evals for LLM systems.

    Verified

  7. DeepSource

    Static analysis SaaS with Autofix, now talking to agents.

    Verified

  8. Galileo

    Evaluation and observability for production AI systems.

    Verified

  9. Giskard

    Open testing and red-teaming for AI applications.

    Verified

  10. Graphite

    Stacked PRs plus AI review, aimed at a faster merge path.

    Verified

  11. Greptile

    AI review that claims to read the whole repo, not the diff alone.

    Verified

  12. Inspect

    UK AISI’s open framework for evaluating language model agents.

    Verified

  13. Langfuse

    Open-source LLM observability you can self-host.

    Verified

  14. LangSmith

    LangChain’s commercial tracing, eval, and prompt hub.

    Verified

  15. LangWatch

    Observability and evaluation platform for LLM products.

    Verified

  16. Mastra

    TypeScript agent framework with workflows and evals attached.

    Verified

  17. Mergify

    Merge queue and automation rules for GitHub.

    Verified

  18. Phoenix

    Arize’s open observability and eval notebook for LLM apps.

    Verified

  19. Promptfoo

    Open evals and red-team for LLM apps and agents.

    Verified

  20. Ragas

    Open metrics for RAG and LLM application quality.

    Verified

  21. Semgrep

    Fast static analysis you can teach, including AI-assisted rules.

    Verified

  22. Trunk

    Unified linters, formatters, and CI flakiness as a merge gate.

    Verified