Harnessing Engineering: Leveraging Codex in an Agent-First World

Discover the innovative approach of using Codex to develop software with zero manually-written code, enhancing engineering efficiency.

Blog cover image
2101050's avatar
2101050
5 views

Harnessing Engineering: Leveraging Codex in an Agent-First World

By Ryan Lopopolo, Member of the Technical Staff

Over the past five months, our team embarked on an ambitious experiment: developing and deploying a software product with zero lines of manually-written code. This product, now in internal beta, is actively used by our team and external alpha testers. It operates, breaks, and gets fixed—all through code written by Codex. We estimate this approach saved us about 90% of the time it would have taken to write the code manually.

Humans steer. Agents execute. This was our guiding principle. We aimed to drastically increase engineering velocity by shifting the focus from writing code to designing environments and specifying intent. This post shares our journey of building a new product with a team of agents—what worked, what didn’t, and how we maximized our most valuable resource: human time and attention.

Starting from Scratch

In August 2025, we made our first commit to an empty repository. Codex, using GPT-5, generated the initial structure, including CI configuration, formatting rules, and application framework. Even the AGENTS.md file, which guides agents on repository work, was Codex-generated. Five months later, the repository boasts nearly a million lines of code, with around 1,500 pull requests merged by a small team of engineers. This translates to an impressive throughput of 3.5 PRs per engineer per day, increasing as the team grew. Importantly, this wasn’t just about output; the product is actively used by hundreds of internal users.

Redefining the Engineer’s Role

Without human-written code, engineering work shifted to focus on systems, scaffolding, and leverage. Initially, progress was slow—not due to Codex’s limitations, but because the environment was underspecified. Our engineers’ primary role became enabling agents to perform useful work. This involved breaking down larger goals into smaller tasks, prompting agents to construct these blocks, and using them to tackle more complex tasks. When issues arose, the solution was never to “try harder.” Instead, engineers asked, “What capability is missing, and how can we make it clear and enforceable for the agent?”

Engineers interacted with the system primarily through prompts: describing tasks, running agents, and allowing them to open pull requests. Codex reviewed its changes, requested additional reviews, and iterated until all feedback was addressed. Humans could review pull requests but weren’t required to. Over time, most review efforts shifted to agent-to-agent interactions.

Enhancing Application Legibility

As code throughput increased, human QA capacity became a bottleneck. To address this, we enhanced agent capabilities by making the application UI, logs, and metrics directly legible to Codex. For instance, we enabled Codex to launch and drive one instance per change, using the Chrome DevTools Protocol to work with DOM snapshots and screenshots. This allowed Codex to reproduce bugs, validate fixes, and understand UI behavior directly.

Diagram: Codex uses Chrome DevTools to validate its work, observing runtime events, applying fixes, and re-running validation until the app is clean.

We applied the same approach to observability tooling. Logs, metrics, and traces are exposed to Codex via a local observability stack, allowing agents to query logs with LogQL and metrics with PromQL. This made prompts like “ensure service startup completes in under 800ms” feasible.

Diagram: Codex uses a full observability stack to query, correlate, and reason, implementing fixes and re-running workloads in a feedback loop.

Repository Knowledge as the System of Record

Managing context is crucial for effective agent work on complex tasks. We learned early on to provide Codex with a map, not a 1,000-page instruction manual. A giant instruction file crowds out essential information, leading to missed constraints or misaligned optimization. Instead, we treat AGENTS.md as a table of contents, with the repository’s knowledge base living in a structured docs/ directory.

Repository Knowledge Store Layout

Text
11AGENTS.md
22ARCHITECTURE.md
3undefinedundefined
4

Recommended Articles

Discover more articles you might find interesting

Implementing LangGraph REST API with FastAPI
Technical Insights

Implementing LangGraph REST API with FastAPI

This guide provides a comprehensive implementation plan for building a LangGraph REST API using FastAPI, covering environment setup, agent definitions, endpoint creation, testing, and deployment.

2101050
Jun 18
149
Read More
DeepSite v2 Practical Guide
Technical Insights

DeepSite v2 Practical Guide

A comprehensive guide to DeepSite v2, covering its features, installation, and advanced workflows.

2101050
Jun 21
110
Read More
Fastify OpenTelemetry: Logging, Metrics, and Tracing in Practice
Technical Insights

Fastify OpenTelemetry: Logging, Metrics, and Tracing in Practice

Learn how to implement logging, metrics, and tracing in Fastify using OpenTelemetry.

2101050
Jul 11
105
Read More
Creating Diverse Logo Designs with Flux Model and ComfyUI
Technical Insights

Creating Diverse Logo Designs with Flux Model and ComfyUI

Learn to leverage the Flux model and ComfyUI for unique logo designs through effective prompts and examples.

2101050
Jan 10
92
Read More
Formatting Dates in TypeScript to UTC
Technical Insights

Formatting Dates in TypeScript to UTC

A guide on how to format dates in TypeScript to the specific format YYYY-MM-DDTHH:mm:ss+00:00.

2101050
Dec 19
82
Read More
Implementing a Custom Chat Model with LangChain
Technical Insights

Implementing a Custom Chat Model with LangChain

This guide provides a comprehensive blueprint for creating a custom chat model by subclassing LangChain's BaseChatModel, including configuration, method overrides, and error handling.

2101050
Jun 17
77
Read More