softwarefactory.build

A directory and guide library for agent-native software factories.

Directory + guides · No pay-to-play

The factory, without the sales deck.

An agent-native software factory takes a signal in one end — an issue, a spec, a customer report — and produces deployed, verified software out the other. Agents execute. Engineers define intent and review results.

108 hand-curated listings, 8 guides, and 4 sourced case studies. Assessments are editorial. Placement is not for sale. The publisher founded Trevize; that conflict is on methodology.


The stack

All listings

Directory

Browse and filter

108 listings, one YAML file each. Every page carries a verifiedAt date and an editorial take. Start with a roundup or a comparison if you already know the job.

  1. Agent Development Kit

    Google’s open kit for building Gemini-centered agents.

    Verified

  2. AgentOps

    Observability aimed at agent runs, not only chat completions.

    Verified

  3. Agno

    Python multi-agent framework (formerly Phidata).

    Verified

  4. Aider

    Git-native CLI agent that commits as it works.

    Verified

  5. Amazon Q Developer

    AWS’s coding assistant and agent, billed into the AWS account.

    Verified

  6. Amp

    Sourcegraph’s agent-native coding product, successor energy to Cody.

    Verified

  7. Augment

    Enterprise coding agent that sells context over huge codebases.

    Verified

  8. Blaxel

    Agent sandbox/compute with a hibernation-shaped billing story.

    Verified

  9. Bolt

    StackBlitz’s in-browser agent that builds apps in WebContainers.

    Verified

  10. Braintrust

    Eval-first platform: datasets, experiments, and logging.

    Verified

  11. Browserbase

    Headless browsers as a service for agents that must click.

    Verified

  12. Claude Code

    Anthropic's coding agent across terminal, IDE, desktop, and browser.

    Verified

  13. Cline

    Open VS Code agent that drives the editor like a human.

    Verified

  14. Cloudflare Agents

    Agents SDK on Durable Objects — this site’s own job runtime.

    Verified

  15. Cloudflare Sandbox

    Isolated code execution on Cloudflare, next to Workers.

    Verified

  16. CodeQL

    GitHub’s semantic code analysis, query-based and CI-native.

    Verified

  17. Coder

    Self-hosted workspaces on your infrastructure, Terraform-defined.

    Verified

  18. CodeRabbit

    High-volume AI pull-request review as a product.

    Verified

  19. Context7

    Up-to-date library docs injected into coding agents.

    Verified

  20. Continue

    Open IDE agents, acquired by Cursor, source still public.

    Verified

  21. CrewAI

    Role-playing multi-agent framework with a commercial control plane.

    Verified

  22. Crush

    Charm’s glamorous terminal coding agent.

    Verified

  23. Cubic

    AI code review product aimed at merging faster with fewer comments.

    Verified

  24. Cursor

    An agent-native IDE with cloud agents that run in isolated VMs.

    Verified

  25. Dagger

    Programmable CI/CD engine agents can call as a tool.

    Verified

  26. Daytona

    Workspaces for agents, with a more persistent posture than E2B.

    Verified

  27. DeepEval

    Pytest-flavored evals for LLM systems.

    Verified

  28. DeepSource

    Static analysis SaaS with Autofix, now talking to agents.

    Verified

  29. Devbox

    Nix-powered reproducible shells for humans and agents.

    Verified

  30. Devin

    Cognition’s cloud software engineer; Windsurf is now Devin Desktop.

    Verified

  31. DSPy

    Stanford’s framework for programming — not prompting — LMs.

    Verified

  32. E2B

    Firecracker microVMs as a sandbox API for AI agents.

    Verified

  33. Ellipsis

    AI reviewer and fixer that can open follow-up PRs.

    Verified

  34. Factory

    Droids across the SDLC, sold as a software factory.

    Verified

  35. Fern

    OpenAPI-to-SDKs and docs, an alternative to Stainless.

    Verified

  36. Firecracker

    AWS’s microVM monitor that sandbox vendors wrap.

    Verified

  37. Fly Machines

    Fast microVMs you API-control, used as agent computers.

    Verified

  38. Galileo

    Evaluation and observability for production AI systems.

    Verified

  39. Gemini CLI

    Google’s open-source terminal agent for Gemini models.

    Verified

  40. Gemini Code Assist

    Google’s IDE coding assistant for Gemini, distinct from Jules.

    Verified

  41. Giskard

    Open testing and red-teaming for AI applications.

    Verified

  42. GitHub Codespaces

    Microsoft-hosted dev VMs that agents and humans already share.

    Verified

  43. GitHub Copilot

    Incumbent coding assistant with a cloud coding agent on GitHub.

    Verified

  44. Gitpod

    Open-source platform for automated, ready-to-code environments.

    Verified

  45. Goose

    Block’s local-first, open-source agent with recipes.

    Verified

  46. Graphite

    Stacked PRs plus AI review, aimed at a faster merge path.

    Verified

  47. Greptile

    AI review that claims to read the whole repo, not the diff alone.

    Verified

  48. Hatchet

    Open task queue aimed at durable, observable background jobs.

    Verified

  49. Haystack

    deepset’s open framework for pipelines, RAG, and agents.

    Verified

  50. Helicone

    Open LLM gateway with logging, caching, and spend controls.

    Verified

  51. Inngest

    Event-driven durable functions for product and agent jobs.

    Verified

  52. Inspect

    UK AISI’s open framework for evaluating language model agents.

    Verified

  53. Jules

    Google’s asynchronous coding agent that opens pull requests.

    Verified

  54. Junie

    JetBrains’ coding agent inside IntelliJ-family IDEs.

    Verified

  55. Kilo

    Open agent across IDE and CLI, plus a zero-markup model gateway.

    Verified

  56. Langfuse

    Open-source LLM observability you can self-host.

    Verified

  57. LangGraph

    Graph runtime for agents, from the LangChain project.

    Verified

  58. LangSmith

    LangChain’s commercial tracing, eval, and prompt hub.

    Verified

  59. LangWatch

    Observability and evaluation platform for LLM products.

    Verified

  60. Letta

    Stateful agents with memory, from the MemGPT lineage.

    Verified

  61. Linear

    Issue tracker that became an intent API for coding agents.

    Verified

  62. LlamaIndex

    Data framework for LLM apps that grew workflows and agents.

    Verified

  63. Logfire

    Pydantic’s observability product, including LLM/agent traces.

    Verified

  64. Lovable

    Prompt-to-full-stack app builder with a hosted computer.

    Verified

  65. Mastra

    TypeScript agent framework with workflows and evals attached.

    Verified

  66. Mergify

    Merge queue and automation rules for GitHub.

    Verified

  67. Microsoft Agent Framework

    Microsoft’s successor surface for AutoGen-style multi-agent apps.

    Verified

  68. Mintlify

    Docs platform that treats agent-readable documentation as a product.

    Verified

  69. Modal

    Serverless containers and GPUs that agents can treat as a computer.

    Verified

  70. n8n

    Open workflow automation that now grows agent nodes.

    Verified

  71. Northflank

    Workloads and sandboxes on your cloud, with an agent-aware pitch.

    Verified

  72. OpenAI Agents SDK

    OpenAI’s official Python/TS SDK for agent loops.

    Verified

  73. OpenAI Codex

    OpenAI’s coding agent across CLI, IDE, and cloud tasks.

    Verified

  74. OpenCode

    Open terminal coding agent; Kilo CLI is a downstream fork.

    Verified

  75. OpenHands

    Open software-agent runtime with a computer you can watch.

    Verified

  76. OpenRouter

    Multi-provider inference router used by many coding agents.

    Verified

  77. Phoenix

    Arize’s open observability and eval notebook for LLM apps.

    Verified

  78. Plandex

    Open terminal agent that plans large diffs in a sandbox.

    Verified

  79. Portkey

    AI gateway: routing, guardrails, and observability in one proxy.

    Verified

  80. Prefect

    Python workflow orchestrator used to schedule agent jobs.

    Verified

  81. Promptfoo

    Open evals and red-team for LLM apps and agents.

    Verified

  82. Pydantic AI

    Typed Python agents from the Pydantic team.

    Verified

  83. Qodo

    Generation and review agents aimed at test-aware pull requests.

    Verified

  84. Ragas

    Open metrics for RAG and LLM application quality.

    Verified

  85. Replit Agent

    Hosted computer plus agent that can build and deploy an app.

    Verified

  86. Roo Code

    Cline fork aimed at modes, automation, and teams.

    Verified

  87. Semgrep

    Fast static analysis you can teach, including AI-assisted rules.

    Verified

  88. smolagents

    Hugging Face’s small, first-class code-agent library.

    Verified

  89. Sourcery

    Refactoring and review assistant with a long IDE history.

    Verified

  90. Spec Kit

    GitHub’s toolkit for spec-driven development with coding agents.

    Verified

  91. SpecStory

    Capture AI IDE sessions so intent does not die in the chat.

    Verified

  92. Stainless

    SDK generation from an OpenAPI spec — a machine-readable intent.

    Verified

  93. Steel

    Open-source browser sandbox API for AI agents.

    Verified

  94. SWE-agent

    Princeton’s agent that made SWE-bench a product category.

    Verified

  95. Tabby

    Self-hosted coding assistant you run like an internal service.

    Verified

  96. Tabnine

    Long-running completion vendor with enterprise model hosting options.

    Verified

  97. Temporal

    Durable workflow engine that can sequence factory steps.

    Verified

  98. Tessl

    Spec-first platform for building software with agents.

    Verified

  99. Trae

    ByteDance’s agentic IDE, free-tier aggressive, closed runtime.

    Verified

  100. Trevize

    Shared cloud workspaces where teams and agents close the loop.

    Verified

  101. Trigger.dev

    Background jobs for TypeScript apps, including long agent runs.

    Verified

  102. Trunk

    Unified linters, formatters, and CI flakiness as a merge gate.

    Verified

  103. v0

    Vercel’s UI generator that ships into a Next.js repo.

    Verified

  104. Vercel AI SDK

    TypeScript SDK for streaming model UIs and tool-calling apps.

    Verified

  105. Vercel Sandbox

    Vercel’s documented sandbox primitive for running untrusted code.

    Verified

  106. Warp

    Open terminal, multi-model agents, and factories in early access.

    Verified

  107. WebContainers

    StackBlitz’s in-browser Node runtime used as an agent sandbox.

    Verified

  108. Zed

    Open Rust editor with first-class agent panels.

    Verified


Guides

All guides
  1. Assembling a self-hosted stack

    How to pick an open execution harness, a sandbox you can inspect, and gates you own — without pretending one vendor is the factory.

    intermediate

  2. Choosing an execution harness

    How to pick the agent that writes the diff — IDE, CLI, or cloud worker — without a bake-off that measures the wrong thing.

    introductory

  3. Designing the intent layer

    How work enters a factory — specs, issues, and pass conditions — so agents have something to close against.

    intermediate

  4. Verification and quality gates

    How to decide a diff is software — tests, review, evals, and merge policy — without trusting a green check the agent painted.

    intermediate

  5. What a software factory is

    The agent-native definition, the older DevSecOps meaning, and the loop that actually has to close.

    introductory

  6. Sandboxing and isolation

    Why the agent needs its own computer, what isolation actually means, and which shortcuts turn a factory into a laptop with extra steps.

    intermediate

  7. Measuring factory output

    What to count when agents write the diffs — and which dashboards are just token spend with a nicer chart.

    advanced

  8. How to read a listing

    What the fields mean, why verifiedAt is load-bearing, and how to treat an editorial take.

    introductory


Case studies

All studies

Production deployments with every scale claim tied to a primary source. Talk-circuit numbers we could not open on the organization’s own site are omitted.

  1. Shopify: River, a Slack-native agent on the Aquifer substrate

    Shopify built River, an AI agent that only operates in public Slack channels. It reads code, runs tests, opens pull requests, queries the warehouse, and inspects production traces. Underneath, Aquifer is the internal platform (durable session, harness, sandbox, gateway, credentials proxy, observability). The May 2026 engineering post is explicit that the 2024 monorepo-and-Nix work was done so agents would have a substrate.

  2. Stripe: Internal minions and AI-assisted pull requests

    Stripe runs internal coding agents it calls minions. In the Sessions 2026 developer keynote, API design lead Michelle Bu described minions that open production pull requests for human review, alongside a broader rise in AI-assisted PR creation and agent traffic to Stripe’s own documentation.

  3. StrongDM: A three-person AI team and non-interactive development

    StrongDM’s AI team published a first-party account of a “Software Factory” in which humans write specifications and scenarios, and agents write and iterate on code without human-authored patches or human code review. The validation story is scenario satisfaction against a Digital Twin Universe (behavioral clones of Okta, Jira, Slack, and Google Workspace APIs), not a conventional green test suite.

  4. Uber: Minion, uReview, and the internal agent stack

    Uber engineers have described Minion as an internal background-agent platform that takes toil (validation-heavy fixes, cleanup, similar interruptions) and returns a pull request for a human. That description comes from talks by Uber staff, not from an uber.com post named Minion. What Uber has published under its own domain is uReview, a GenAI review system that comments on diffs, plus adjacent tools such as FixrLeak. This page keeps those sources separate.


Compare

Meaningful pairs only. Not a generated cross-product.

All comparisons

Roundups

Inclusion is a filter on structured fields, not a rank.


Build log

This site was assembled by the same kind of factory it describes. The log records the pipeline, what broke, and what agents got wrong.

Finishing the brief — pages, sources, and a hundred listings

All entries

How this is run

Listings live as YAML in Git. Assessments are opinionated and sourced. We do not sell placement. The publisher founded Trevize; that conflict is on the methodology page, not buried in a footer.

Propose a listing