softwarefactory.build

A directory and guide library for agent-native software factories.

Layer / What must pass

Verification

Tests, reviews, hooks, and merge gates that decide whether output is software or just a plausible patch.


  1. Amazon Q Developer

    AWS’s coding assistant and agent, billed into the AWS account.

    Verified

  2. Braintrust

    Eval-first platform: datasets, experiments, and logging.

    Verified

  3. Claude Code

    Anthropic's coding agent across terminal, IDE, desktop, and browser.

    Verified

  4. CodeQL

    GitHub’s semantic code analysis, query-based and CI-native.

    Verified

  5. CodeRabbit

    High-volume AI pull-request review as a product.

    Verified

  6. Cubic

    AI code review product aimed at merging faster with fewer comments.

    Verified

  7. Cursor

    An agent-native IDE with cloud agents that run in isolated VMs.

    Verified

  8. Dagger

    Programmable CI/CD engine agents can call as a tool.

    Verified

  9. DeepEval

    Pytest-flavored evals for LLM systems.

    Verified

  10. DeepSource

    Static analysis SaaS with Autofix, now talking to agents.

    Verified

  11. Devin

    Cognition’s cloud software engineer; Windsurf is now Devin Desktop.

    Verified

  12. Ellipsis

    AI reviewer and fixer that can open follow-up PRs.

    Verified

  13. Factory

    Droids across the SDLC, sold as a software factory.

    Verified

  14. Galileo

    Evaluation and observability for production AI systems.

    Verified

  15. Giskard

    Open testing and red-teaming for AI applications.

    Verified

  16. GitHub Copilot

    Incumbent coding assistant with a cloud coding agent on GitHub.

    Verified

  17. Graphite

    Stacked PRs plus AI review, aimed at a faster merge path.

    Verified

  18. Greptile

    AI review that claims to read the whole repo, not the diff alone.

    Verified

  19. Inspect

    UK AISI’s open framework for evaluating language model agents.

    Verified

  20. Langfuse

    Open-source LLM observability you can self-host.

    Verified

  21. LangSmith

    LangChain’s commercial tracing, eval, and prompt hub.

    Verified

  22. LangWatch

    Observability and evaluation platform for LLM products.

    Verified

  23. Mastra

    TypeScript agent framework with workflows and evals attached.

    Verified

  24. Mergify

    Merge queue and automation rules for GitHub.

    Verified

  25. Phoenix

    Arize’s open observability and eval notebook for LLM apps.

    Verified

  26. Promptfoo

    Open evals and red-team for LLM apps and agents.

    Verified

  27. Qodo

    Generation and review agents aimed at test-aware pull requests.

    Verified

  28. Ragas

    Open metrics for RAG and LLM application quality.

    Verified

  29. Semgrep

    Fast static analysis you can teach, including AI-assisted rules.

    Verified

  30. Sourcery

    Refactoring and review assistant with a long IDE history.

    Verified

  31. SWE-agent

    Princeton’s agent that made SWE-bench a product category.

    Verified

  32. Trevize

    Shared cloud workspaces where teams and agents close the loop.

    Verified

  33. Trunk

    Unified linters, formatters, and CI flakiness as a merge gate.

    Verified