Building a Platform Stateless Runner with LangGraph, Redis, and FastAPI
1. Introduction: Why Stateless Runners?
Overview
In today’s distributed computing ecosystem, the adoption of stateless application patterns has become crucial for scalability, reliability, and platform flexibility. Stateless runners—compute components that do not depend on local or in-memory state—enable all critical state to persist externally, typically in a dedicated database, cache, or queue system. This allows you to scale horizontally, recover seamlessly from failures, and support cloud-native, microservice-oriented workflows.
When do you need this?
- Batch and pipeline orchestration: Workloads are split across short-lived tasks.
- Asynchronous API tasks: Users submit requests and poll for results.
- Large-scale LLM or data processing: Jobs can be split, resumed, and recovered without loss.
Stack Overview:
- LangGraph manages the workflow logic as a directed acyclic graph, chaining together multiple steps including validations, LLM calls, and summaries.
- Redis serves as the high-performance, in-memory key/value store, ensuring that all progress, errors, and final results are persisted externally.
- FastAPI exposes APIs for users to submit jobs and query results, acting as a stateless gateway.
Architectural Diagram (Mermaid)
Mermaid
Sample Request Flow Example
- User sends a POST to
/jobswith input. - FastAPI assigns a
job_idand stores an initial job state in Redis. - FastAPI kicks off the LangGraph workflow with this state.
- Each LangGraph step updates the job state in Redis.
- The user polls GET
/jobs/{job_id}, and FastAPI fetches the latest state from Redis.
References
2. Environment Setup and Prerequisites
2.1. Project Structure
Text
2.2. Python Virtual Environment & Package Setup
Sh
In requirements.txt:
Text
Sh
2.3. Redis (Docker-based Setup)
Sh
Confirm it works:
Sh
2.4. Minimal FastAPI + Redis Test
app/main.py:
Python
Run:
Sh
3. Deep Dive: Stateless Runner Implementation
3.1. Workflow State and LangGraph Setup
app/models.py
Python
app/workflow.py
Python
3.2. Redis Integration
app/redis_client.py:
Python
3.3. FastAPI API Layer
app/main.py:
Python
Run & test as above.
4. End-to-End Practical Example & Case Study
4.1. Workflow & API Integration
More advanced workflow (with retries, error handling) as in Section 4 above—see detailed code samples in the previous message.
4.2. Testing
Bash
4.3. Troubleshooting Table
| Symptom | Cause | Fix |
|---|---|---|
| Always "pending" | Logic error | Add debug logs |
| "Job not found" | Wrong key/exp | Validate and check TTL |
| Data lost on fail | Redis not persist | Enable disk volume |
See Section 4 above for full example code and diagrams.
5. Further Considerations, Troubleshooting & Extensibility
5.1. Troubleshooting Checklist
- Jobs persist after crash? If not, check Redis persistence.
- API unreachable? Docker/Compose service config, logs.
- Unexpected errors? Add deeper logging in workflow and Redis client wrappers.
5.2. Security: API & Redis
API Key Security:
Python
Redis password:
- Add
requirepassin configuration. - Use in the Python Redis client.
5.3. Production Deployment
docker-compose.yml:
Yaml
Kubernetes (section above for manifests and patterns).
5.4. Observability & Monitoring
- Use Prometheus + Grafana.
- Add
/metricsendpoint with Prometheus FastAPI Instrumentator - Monitor Redis with Redis Exporter
5.5. Extensibility
- Introduce Celery/advanced queues for heavier workloads.
- Swap Redis for DB/other key-value stores.
- Register multiple workflow types with LangGraph.






