Implementation Blueprint: Required Overrides and Configuration
This section lays out the minimal interface contract you must fulfill to wrap any external chat API under LangChain’s BaseChatModel abstraction.
1.1 Core Class Signature
Your custom model subclass must inherit from langchain_core.language_models.BaseChatModel. For example, in v0.3 of the docs:
Python
- Pydantic fields become the constructor args (e.g.,
MyCustomChatModel(model="x", api_key="…")). - Use
Field(alias="model")if your API expects a different key.
1.2 Mandatory Method Overrides
At minimum, override:
-
_generate(self, messages, stop=None, run_manager=None, **kwargs) -> ChatResult -
_agenerate(self, messages, stop=None, run_manager=None, **kwargs) -> ChatResult -
_stream(self, messages, stop=None, run_manager=None, **kwargs) -> Iterator[ChatGenerationChunk](only if your API supports streaming) -
_identifying_params(self) -> Dict[str, Any]
The v0.3 guide’s ChatParrotLink example shows a skeleton:
Python
1.3 Callback Integration Checklist
run_manager.on_llm_startbefore sending promptrun_manager.on_llm_streamfor each token/chunk (in streaming)run_manager.on_llm_endafter final response- For async: use
self.astream_eventsto emit events
1.4 Packaging Responses
- Sync: return
ChatResult(generations=[ChatGeneration(… )]) - Async: same return type; LangChain adapts
ainvoke - Streaming: iterate and yield
ChatGenerationChunkobjects, containing partialAIMessageChunk
1.5 Identifying Params
Ensure _identifying_params returns a dict of all config fields that define which underlying model variant you’re calling (e.g., model name, temperature). LangChain uses this for hashing and caching.
2. Subclassing BaseChatModel: Sync & Async Methods
Now we build the custom wrapper step by step, focusing on synchronous and asynchronous path implementations.
2.1 Defining Configuration Fields
Python
- Type hints and
Fielddescriptions guarantee correct initialization and auto-documentation in LangChain’s CLI or notebooks.
2.2 Implementing _generate
Python
Notes:
- Collect token counts from
usageif available, else compute via string length. - Include
**self._identifying_params()in callback events.
2.3 Implementing _agenerate
Python
2.4 _identifying_params
Python
This ensures LangChain can key cache and telemetry by unique model settings.
3. Streaming & Callback Integration
When your API offers streaming (chunked) responses, you can surface partial tokens to LangChain’s streaming interface.
3.1 Streaming Method Signature
Python
3.2 Async Streaming
With httpx.AsyncClient:
Python
3.3 Astream Events API
LangChain’s newer astream_events lets you attach a single event generator:
Python
Use this if you want unified async event handling and avoid duplicating callback calls.
4. Testing, Batching, Error Handling, Deployment
A production-ready custom chat model also includes:
4.1 Batch Support (Threadpool)
LangChain auto-wraps your sync _generate into batch calls via a threadpool by default. To manually customize:
Python
Or rely on inherited .batch() behavior.
4.2 Unit & Integration Tests
- Sync invoke:
Python
- Async ainvoke:
Python
- Streaming: iterate and assemble chunks.
4.3 Error Handling & Retries
- Use exponential backoff:
Python
- Catch HTTP 5xx and 429 codes specially.
- Raise on final failure.
4.4 Deployment & Configuration
- Read secrets from environment:
Python






