Harnessing Engineering: Leveraging Codex in an Agent-First World
By Ryan Lopopolo, Member of the Technical Staff
Over the past five months, our team embarked on an ambitious experiment: developing and deploying a software product with zero lines of manually-written code. This product, now in internal beta, is actively used by our team and external alpha testers. It operates, breaks, and gets fixed—all through code written by Codex. We estimate this approach saved us about 90% of the time it would have taken to write the code manually.
Humans steer. Agents execute. This was our guiding principle. We aimed to drastically increase engineering velocity by shifting the focus from writing code to designing environments and specifying intent. This post shares our journey of building a new product with a team of agents—what worked, what didn’t, and how we maximized our most valuable resource: human time and attention.
Starting from Scratch
In August 2025, we made our first commit to an empty repository. Codex, using GPT-5, generated the initial structure, including CI configuration, formatting rules, and application framework. Even the AGENTS.md file, which guides agents on repository work, was Codex-generated. Five months later, the repository boasts nearly a million lines of code, with around 1,500 pull requests merged by a small team of engineers. This translates to an impressive throughput of 3.5 PRs per engineer per day, increasing as the team grew. Importantly, this wasn’t just about output; the product is actively used by hundreds of internal users.
Redefining the Engineer’s Role
Without human-written code, engineering work shifted to focus on systems, scaffolding, and leverage. Initially, progress was slow—not due to Codex’s limitations, but because the environment was underspecified. Our engineers’ primary role became enabling agents to perform useful work. This involved breaking down larger goals into smaller tasks, prompting agents to construct these blocks, and using them to tackle more complex tasks. When issues arose, the solution was never to “try harder.” Instead, engineers asked, “What capability is missing, and how can we make it clear and enforceable for the agent?”
Engineers interacted with the system primarily through prompts: describing tasks, running agents, and allowing them to open pull requests. Codex reviewed its changes, requested additional reviews, and iterated until all feedback was addressed. Humans could review pull requests but weren’t required to. Over time, most review efforts shifted to agent-to-agent interactions.
Enhancing Application Legibility
As code throughput increased, human QA capacity became a bottleneck. To address this, we enhanced agent capabilities by making the application UI, logs, and metrics directly legible to Codex. For instance, we enabled Codex to launch and drive one instance per change, using the Chrome DevTools Protocol to work with DOM snapshots and screenshots. This allowed Codex to reproduce bugs, validate fixes, and understand UI behavior directly.

We applied the same approach to observability tooling. Logs, metrics, and traces are exposed to Codex via a local observability stack, allowing agents to query logs with LogQL and metrics with PromQL. This made prompts like “ensure service startup completes in under 800ms” feasible.
Repository Knowledge as the System of Record
Managing context is crucial for effective agent work on complex tasks. We learned early on to provide Codex with a map, not a 1,000-page instruction manual. A giant instruction file crowds out essential information, leading to missed constraints or misaligned optimization. Instead, we treat AGENTS.md as a table of contents, with the repository’s knowledge base living in a structured docs/ directory.
Repository Knowledge Store Layout
Text






