Developer Guide: Middleware/api/
1. Overview
This directory contains the primary entry point for the WilmerAI application: a Flask web server that exposes all public-facing API endpoints. Its fundamental role is to act as a robust compatibility and translation layer. It accepts requests conforming to popular external schemas (like OpenAI and Ollama), transforms the data into a standardized internal format, and dispatches the request to the central workflow engine for processing.
The architecture is designed to be modular and extensible. The API logic is broken down into an orchestrator (ApiServer), a business logic gateway (workflow_gateway), and a series of self-contained handlers for specific API schemas. A key component of this architecture is the ResponseBuilderService, which centralizes all logic for constructing API-specific JSON responses, ensuring that outgoing data matches the schema expected by the client.
This separation of concerns makes the system cleaner, more maintainable, and significantly easier to extend with new endpoints or entire API compatibility layers.
2. Architectural Flow
A request's journey through the api package follows a well-defined sequence. The primary difference in the flow
depends on whether the client requests a streaming or non-streaming response.
Ingestion: An external client sends an HTTP request to a specific URL (e.g., /v1/chat/completions).
Routing: The Flask application, orchestrated by ApiServer at startup, routes the request to the
appropriate MethodView class (e.g., ChatCompletionsAPI) located within a registered handler file (
e.g., openai_api_handler.py).
Method Invocation: The corresponding HTTP method within the class is called (typically post()). The method
immediately performs two key setup actions:
- It generates a unique
request_id and stores it in Flask's g context object. This ID is used throughout the
stack to enable request cancellation.
- It sets a global
API_TYPE variable (e.g., openaichatcompletion) so downstream components know which response
format to use.
Pre-Processing & Validation: The method performs initial actions, such as applying decorators (
e.g., @handle_openwebui_tool_check) and validating the incoming JSON payload.
Data Transformation: The endpoint's logic transforms the client-specific payload into a standardized
internal messages list. This includes parsing different structures, handling multimedia content like base64 images,
and applying configuration-based transformations.
Engine Handoff: The request_id, transformed messages list, and a boolean stream flag are passed to
the handle_user_prompt() function in workflow_gateway.py. This is the critical handoff point where control is
passed to the internal workflow engine.
Response Handling: The handler waits for a result from the workflow engine. The handling depends on the stream
flag:
- Streaming: If
stream is true, handle_user_prompt() returns a Python generator. The handler immediately
returns a Flask Response object with the appropriate mimetype (e.g., 'text/event-stream'). The handler is
specifically designed to work with production WSGI servers like Eventlet to detect client disconnects and
trigger cancellation. Deep within the streaming logic, api_helpers.build_response_json() is called for each
token, which in turn uses the ResponseBuilderService to construct a well formatted JSON chunk string for
the specific API type.
- Non-Streaming: If
stream is false, handle_user_prompt() returns a single, complete string of text. The
handler then calls the appropriate method directly on the ResponseBuilderService (
e.g., response_builder.build_openai_chat_completion_response(text)) to construct the final, complete JSON
object. This object is then returned to the client using Flask's jsonify.
3. Directory & File Breakdown
Middleware/api/
app.py
- Responsibility: Instantiates the global Flask
app object.
- Rationale: By placing the app instance in its own file, other modules can import it without creating circular
dependencies.
api_server.py
- Responsibility: Discovers all API handlers, registers their routes with the Flask app, and exposes the app for a
WSGI server to run. It is the main orchestrator for the API layer.
- Key Components:
ApiServer class:
_discover_and_register_handlers(): Automatically scans the Middleware/api/handlers/impl/ directory for
Python files. It imports them, finds any classes that inherit from BaseApiHandler, and calls
their register_routes() method, making the system plug-and-play.
run(): Starts the Flask development server (for debugging only).
workflow_gateway.py
- Responsibility: Acts as the single, standardized bridge between the API request handlers and the backend workflow
engine.
- Key Components:
handle_user_prompt(): The crucial function that every data-handling endpoint calls. It takes the request_id,
the standardized messages list, and a stream flag, and dispatches the request to the correct backend service.
@handle_openwebui_tool_check(): A decorator that inspects incoming requests for a specific tool-selection prompt
from clients like OpenWebUI and returns a valid, empty tool-call response to allow the client to proceed
gracefully.
api_helpers.py
- Responsibility: Provides helper functions used across the API layer, primarily for handling streaming data and
workflow override management.
- Key Functions:
build_response_json(): Constructs a single streaming JSON chunk. It acts as a dispatcher, calling the
appropriate chunk-building method on the ResponseBuilderService based on the globally set API_TYPE.
extract_text_from_chunk(): Parses an incoming data chunk from an LLM backend and extracts the text content,
handling multiple formats (SSE, plain JSON).
sse_format(): Formats a string into the correct Server-Sent Event (SSE) structure.
get_model_name(): Returns the model identifier for API responses. When a workflow override is active, returns
username:workflow format; otherwise returns just the username.
parse_model_field(): Parses the model field from incoming API requests and extracts a workflow name if one is
specified. Supports formats like username:workflow, workflow, and username:workflow:latest.
set_workflow_override(): Sets the global workflow override from the model field. Called at the start of request
processing.
clear_workflow_override(): Clears the workflow override. Called at the end of request processing.
get_active_workflow_override(): Returns the currently active workflow override, if any.
Middleware/services/
response_builder_service.py
- Responsibility: To act as the single source of truth for all API response schemas. This service centralizes
the logic for creating JSON payloads, decoupling the response format from the handler and helper logic.
- Key Components:
ResponseBuilderService class:
- Non-Streaming Methods: Contains methods like
build_openai_chat_completion_response()
and build_ollama_generate_response() that accept the final generated text and return a complete,
schema-compliant dictionary.
- Streaming Chunk Methods: Contains methods like
build_openai_chat_completion_chunk()
and build_ollama_chat_chunk() that create a single, schema-compliant JSON chunk for a streaming response.
- Models List Methods:
build_openai_models_response() and build_ollama_tags_response() return lists of
available workflows from the _shared folder. Each workflow is presented in username:workflow format,
allowing front-end applications to select specific workflows via their model dropdown.
Middleware/api/handlers/
base/base_api_handler.py
- Responsibility: Defines the abstract interface (
BaseApiHandler) that all API handlers must implement.
- Key Components:
@abstractmethod register_routes(): The contract method. Every concrete handler must implement this to register
its URL rules with the Flask app.
impl/openai_api_handler.py
- Responsibility: Implements all endpoints that conform to the OpenAI API specification.
- Key
MethodView Classes:
ModelsAPI: Handles /v1/models.
CompletionsAPI: Handles the legacy /v1/completions.
ChatCompletionsAPI: Handles the standard /v1/chat/completions.
- Interactions: Generates a
request_id for each call. Uses workflow_gateway.handle_user_prompt() to process
requests. Relies on eventlet and client disconnection to trigger the CancellationService.
impl/ollama_api_handler.py
- Responsibility: Implements all endpoints that conform to the Ollama API specification, plus extensions for
cancellation.
- Key
MethodView Classes:
GenerateAPI: Handles POST /api/generate.
ApiChatAPI: Handles POST /api/chat.
TagsAPI: Handles /api/tags.
VersionAPI: Handles /api/version.
CancelChatAPI / CancelGenerateAPI: Handle DELETE requests to /api/chat and /api/generate,
respectively. This is a WilmerAI-specific extension that allows clients to explicitly cancel a running request by
providing its request_id.
- Interactions: Generates a
request_id and includes it in responses to the client.
Uses workflow_gateway.handle_user_prompt() to process requests. Streaming responses include disconnect detection
that triggers cancellation.
4. Workflow Selection via Model Field
WilmerAI allows front-end applications to select specific workflows by using the model field in API requests. This
feature enables users to switch between different workflows without changing their configuration.
How It Works
Models List Endpoints: The /v1/models (OpenAI) and /api/tags (Ollama) endpoints return a list of available
workflows from Public/Configs/Workflows/_shared/. Each workflow is presented in username:workflow format.
Request Processing: When a request includes a model field containing a workflow name, the API handler:
- Calls
api_helpers.set_workflow_override() to parse and store the workflow name
- The workflow gateway checks for an active override before using normal routing
- If a valid workflow is found in
_shared/, it's executed directly
Response Format: Responses include the model in username:workflow format when a workflow override is active,
allowing front-ends to maintain the selected workflow across requests.
Model Field Format
The parse_model_field() function in api_helpers.py handles these formats:
| Format |
Example |
Result |
username:workflow |
chris:openwebui-coding |
Extracts openwebui-coding |
workflow |
openwebui-coding |
Uses openwebui-coding if it exists in _shared/ |
username:workflow:latest |
chris:openwebui-coding:latest |
Strips :latest, extracts openwebui-coding |
| Non-matching |
gpt-4 |
Returns None, uses normal routing |
Folder Structure
Public/Configs/Workflows/
├── _shared/
│ ├── openwebui-coding/ # Listed by models endpoint as folder name
│ │ └── _DefaultWorkflow.json # Workflow loaded when folder is selected
│ ├── openwebui-general/
│ │ └── _DefaultWorkflow.json
│ └── openwebui-task/
│ └── _DefaultWorkflow.json
├── chris/ # Default user folder
│ └── ...
Request Flow with Workflow Override
- Client queries
/v1/models → receives list of username:workflow entries
- Client sends request with
"model": "chris:openwebui-coding"
- Handler calls
api_helpers.set_workflow_override("chris:openwebui-coding")
workflow_gateway.handle_user_prompt() detects the override
- Workflow is loaded from
_shared/openwebui-coding/_DefaultWorkflow.json
- Response includes
"model": "chris:openwebui-coding"
- Handler calls
api_helpers.clear_workflow_override() in finally block
Key Files
api_helpers.py: Contains parse_model_field(), set_workflow_override(), clear_workflow_override()
instance_global_variables.py: Contains WORKFLOW_OVERRIDE global variable
workflow_gateway.py: Checks for workflow override before normal routing
config_utils.py: Contains get_shared_workflows_folder(), get_available_shared_workflows(),
workflow_exists_in_shared_folder()
5. How to Extend
The architecture makes it easy to add new functionality. The ApiServer automatically discovers new handlers, so in
most cases, you only need to add new files without modifying existing ones.
Example 2: Add Support for a New API Type (e.g., Anthropic)
This more advanced example shows how to add compatibility for an entirely new API schema, supporting both streaming and
non-streaming.
Update ResponseBuilderService: Add methods to build both the final response and the streaming chunks for the
new API type.
Update api_helpers for Streaming: Wire the new streaming chunk builder into the streaming dispatcher
in build_response_json.
Create the New Handler File: Create the file that will define the endpoints and translation logic for the new
API.
- Create
Middleware/api/handlers/impl/anthropic_api_handler.py.
- Implement the full logic. This includes generating a
request_id, parsing the incoming request, setting
the API_TYPE, calling the gateway, and using the correct response builder methods for both streaming and
non-streaming cases.
# In Middleware/api/handlers/impl/anthropic_api_handler.py
import uuid
from flask import request, jsonify, Response, g
from flask.views import MethodView
from Middleware.api.app import app
from Middleware.api.handlers.base.base_api_handler import BaseApiHandler
from Middleware.api.workflow_gateway import handle_user_prompt
from Middleware.common import instance_global_variables
from Middleware.services.response_builder_service import ResponseBuilderService
response_builder = ResponseBuilderService()
class AnthropicChatAPI(MethodView):
@staticmethod
def post() -> Response:
# 1. Generate request ID for cancellation tracking
request_id = str(uuid.uuid4())
g.current_request_id = request_id
# 2. Set the API type so downstream services know the format
instance_global_variables.API_TYPE = "anthropicchat"
data = request.get_json()
# 3. Transform the Anthropic request into our internal format
messages = [{"role": m["role"], "content": m["content"][0]["text"]} for m in data.get("messages", [])]
stream = data.get("stream", True)
# 4. Handoff to the workflow engine
llm_response = handle_user_prompt(request_id, messages, stream=stream)
if stream:
return Response(llm_response, mimetype='text/event-stream')
else:
response_payload = response_builder.build_anthropic_response(llm_response)
return jsonify(response_payload)
class AnthropicApiHandler(BaseApiHandler):
def register_routes(self, app_instance: app):
app_instance.add_url_rule(
'/v1/anthropic/messages',
view_func=AnthropicChatAPI.as_view('anthropic_chat_api')
)
1---2name: 1-overview3description: This directory contains the primary entry point for the WilmerAI application: a Flask web server that exposes all public-facing API endpoints.4---5-----67### **Developer Guide: `Middleware/api/`**89## 1\. Overview1011This directory contains the primary entry point for the WilmerAI application: a Flask web server that exposes all public-facing API endpoints. Its fundamental role is to act as a robust **compatibility and translation layer**. It accepts requests conforming to popular external schemas (like OpenAI and Ollama), transforms the data into a standardized internal format, and dispatches the request to the central workflow engine for processing.1213The architecture is designed to be modular and extensible. The API logic is broken down into an orchestrator (`ApiServer`), a business logic gateway (`workflow_gateway`), and a series of self-contained **handlers** for specific API schemas. A key component of this architecture is the `ResponseBuilderService`, which centralizes all logic for constructing API-specific JSON responses, ensuring that outgoing data matches the schema expected by the client.1415This separation of concerns makes the system cleaner, more maintainable, and significantly easier to extend with new endpoints or entire API compatibility layers.1617-----1819## 2\. Architectural Flow2021A request's journey through the `api` package follows a well-defined sequence. The primary difference in the flow22depends on whether the client requests a streaming or non-streaming response.23241. **Ingestion**: An external client sends an HTTP request to a specific URL (e.g., `/v1/chat/completions`).25262. **Routing**: The Flask application, orchestrated by `ApiServer` at startup, routes the request to the27 appropriate `MethodView` class (e.g., `ChatCompletionsAPI`) located within a registered handler file (28 e.g., `openai_api_handler.py`).29303. **Method Invocation**: The corresponding HTTP method within the class is called (typically `post()`). The method31 immediately performs two key setup actions:3233 * It generates a unique **`request_id`** and stores it in Flask's `g` context object. This ID is used throughout the34 stack to enable request cancellation.35 * It sets a global `API_TYPE` variable (e.g., `openaichatcompletion`) so downstream components know which response36 format to use.37384. **Pre-Processing & Validation**: The method performs initial actions, such as applying decorators (39 e.g., `@handle_openwebui_tool_check`) and validating the incoming JSON payload.40415. **Data Transformation**: The endpoint's logic transforms the client-specific payload into a standardized42 internal `messages` list. This includes parsing different structures, handling multimedia content like base64 images,43 and applying configuration-based transformations.44456. **Engine Handoff**: The **`request_id`**, transformed `messages` list, and a boolean `stream` flag are passed to46 the `handle_user_prompt()` function in `workflow_gateway.py`. This is the critical handoff point where control is47 passed to the internal workflow engine.48497. **Response Handling**: The handler waits for a result from the workflow engine. The handling depends on the `stream`50 flag:5152 * **Streaming**: If `stream` is true, `handle_user_prompt()` returns a Python generator. The handler immediately53 returns a Flask `Response` object with the appropriate `mimetype` (e.g., `'text/event-stream'`). The handler is54 specifically designed to work with production WSGI servers like `Eventlet` to detect client disconnects and55 trigger cancellation. Deep within the streaming logic, `api_helpers.build_response_json()` is called for each56 token, which in turn uses the `ResponseBuilderService` to construct a well formatted JSON chunk string for57 the specific API type.58 * **Non-Streaming**: If `stream` is false, `handle_user_prompt()` returns a single, complete string of text. The59 handler then calls the appropriate method **directly** on the `ResponseBuilderService` (60 e.g., `response_builder.build_openai_chat_completion_response(text)`) to construct the final, complete JSON61 object. This object is then returned to the client using Flask's `jsonify`.6263-----6465## 3\. Directory & File Breakdown6667### `Middleware/api/`6869#### `app.py`7071* **Responsibility**: Instantiates the global Flask `app` object.72* **Rationale**: By placing the app instance in its own file, other modules can import it without creating circular73 dependencies.7475#### `api_server.py`7677* **Responsibility**: Discovers all API handlers, registers their routes with the Flask app, and exposes the app for a78 WSGI server to run. It is the main orchestrator for the API layer.79* **Key Components**:80 * `ApiServer` class:81 * `_discover_and_register_handlers()`: Automatically scans the `Middleware/api/handlers/impl/` directory for82 Python files. It imports them, finds any classes that inherit from `BaseApiHandler`, and calls83 their `register_routes()` method, making the system plug-and-play.84 * `run()`: Starts the Flask development server (for debugging only).8586#### `workflow_gateway.py`8788* **Responsibility**: Acts as the single, standardized bridge between the API request handlers and the backend workflow89 engine.90* **Key Components**:91 * `handle_user_prompt()`: The crucial function that every data-handling endpoint calls. It takes the `request_id`,92 the standardized `messages` list, and a `stream` flag, and dispatches the request to the correct backend service.93 * `@handle_openwebui_tool_check()`: A decorator that inspects incoming requests for a specific tool-selection prompt94 from clients like OpenWebUI and returns a valid, empty tool-call response to allow the client to proceed95 gracefully.9697#### `api_helpers.py`9899* **Responsibility**: Provides helper functions used across the API layer, primarily for handling streaming data and100 workflow override management.101* **Key Functions**:102 * `build_response_json()`: Constructs a **single streaming JSON chunk**. It acts as a dispatcher, calling the103 appropriate chunk-building method on the `ResponseBuilderService` based on the globally set `API_TYPE`.104 * `extract_text_from_chunk()`: Parses an incoming data chunk from an LLM backend and extracts the text content,105 handling multiple formats (SSE, plain JSON).106 * `sse_format()`: Formats a string into the correct Server-Sent Event (SSE) structure.107 * `get_model_name()`: Returns the model identifier for API responses. When a workflow override is active, returns108 `username:workflow` format; otherwise returns just the username.109 * `parse_model_field()`: Parses the model field from incoming API requests and extracts a workflow name if one is110 specified. Supports formats like `username:workflow`, `workflow`, and `username:workflow:latest`.111 * `set_workflow_override()`: Sets the global workflow override from the model field. Called at the start of request112 processing.113 * `clear_workflow_override()`: Clears the workflow override. Called at the end of request processing.114 * `get_active_workflow_override()`: Returns the currently active workflow override, if any.115116### `Middleware/services/`117118#### `response_builder_service.py`119120* **Responsibility**: To act as the **single source of truth for all API response schemas**. This service centralizes121 the logic for creating JSON payloads, decoupling the response format from the handler and helper logic.122* **Key Components**:123 * `ResponseBuilderService` class:124 * **Non-Streaming Methods**: Contains methods like `build_openai_chat_completion_response()`125 and `build_ollama_generate_response()` that accept the final generated text and return a complete,126 schema-compliant dictionary.127 * **Streaming Chunk Methods**: Contains methods like `build_openai_chat_completion_chunk()`128 and `build_ollama_chat_chunk()` that create a single, schema-compliant JSON chunk for a streaming response.129 * **Models List Methods**: `build_openai_models_response()` and `build_ollama_tags_response()` return lists of130 available workflows from the `_shared` folder. Each workflow is presented in `username:workflow` format,131 allowing front-end applications to select specific workflows via their model dropdown.132133### `Middleware/api/handlers/`134135#### `base/base_api_handler.py`136137* **Responsibility**: Defines the abstract interface (`BaseApiHandler`) that all API handlers must implement.138* **Key Components**:139 * `@abstractmethod register_routes()`: The contract method. Every concrete handler must implement this to register140 its URL rules with the Flask app.141142#### `impl/openai_api_handler.py`143144* **Responsibility**: Implements all endpoints that conform to the OpenAI API specification.145* **Key `MethodView` Classes**:146 * `ModelsAPI`: Handles `/v1/models`.147 * `CompletionsAPI`: Handles the legacy `/v1/completions`.148 * `ChatCompletionsAPI`: Handles the standard `/v1/chat/completions`.149* **Interactions**: Generates a `request_id` for each call. Uses `workflow_gateway.handle_user_prompt()` to process150 requests. Relies on `eventlet` and client disconnection to trigger the `CancellationService`.151152#### `impl/ollama_api_handler.py`153154* **Responsibility**: Implements all endpoints that conform to the Ollama API specification, plus extensions for155 cancellation.156* **Key `MethodView` Classes**:157 * `GenerateAPI`: Handles `POST /api/generate`.158 * `ApiChatAPI`: Handles `POST /api/chat`.159 * `TagsAPI`: Handles `/api/tags`.160 * `VersionAPI`: Handles `/api/version`.161 * **`CancelChatAPI` / `CancelGenerateAPI`**: Handle `DELETE` requests to `/api/chat` and `/api/generate`,162 respectively. This is a WilmerAI-specific extension that allows clients to explicitly cancel a running request by163 providing its `request_id`.164* **Interactions**: Generates a `request_id` and includes it in responses to the client.165 Uses `workflow_gateway.handle_user_prompt()` to process requests. Streaming responses include disconnect detection166 that triggers cancellation.167168-----169170## 4\. Workflow Selection via Model Field171172WilmerAI allows front-end applications to select specific workflows by using the model field in API requests. This173feature enables users to switch between different workflows without changing their configuration.174175### How It Works1761771. **Models List Endpoints**: The `/v1/models` (OpenAI) and `/api/tags` (Ollama) endpoints return a list of available178 workflows from `Public/Configs/Workflows/_shared/`. Each workflow is presented in `username:workflow` format.1791802. **Request Processing**: When a request includes a model field containing a workflow name, the API handler:181 * Calls `api_helpers.set_workflow_override()` to parse and store the workflow name182 * The workflow gateway checks for an active override before using normal routing183 * If a valid workflow is found in `_shared/`, it's executed directly1841853. **Response Format**: Responses include the model in `username:workflow` format when a workflow override is active,186 allowing front-ends to maintain the selected workflow across requests.187188### Model Field Format189190The `parse_model_field()` function in `api_helpers.py` handles these formats:191192| Format | Example | Result |193|--------|---------|--------|194| `username:workflow` | `chris:openwebui-coding` | Extracts `openwebui-coding` |195| `workflow` | `openwebui-coding` | Uses `openwebui-coding` if it exists in `_shared/` |196| `username:workflow:latest` | `chris:openwebui-coding:latest` | Strips `:latest`, extracts `openwebui-coding` |197| Non-matching | `gpt-4` | Returns `None`, uses normal routing |198199### Folder Structure200201```202Public/Configs/Workflows/203├── _shared/204│ ├── openwebui-coding/ # Listed by models endpoint as folder name205│ │ └── _DefaultWorkflow.json # Workflow loaded when folder is selected206│ ├── openwebui-general/207│ │ └── _DefaultWorkflow.json208│ └── openwebui-task/209│ └── _DefaultWorkflow.json210├── chris/ # Default user folder211│ └── ...212```213214### Request Flow with Workflow Override2152161. Client queries `/v1/models` → receives list of `username:workflow` entries2172. Client sends request with `"model": "chris:openwebui-coding"`2183. Handler calls `api_helpers.set_workflow_override("chris:openwebui-coding")`2194. `workflow_gateway.handle_user_prompt()` detects the override2205. Workflow is loaded from `_shared/openwebui-coding/_DefaultWorkflow.json`2216. Response includes `"model": "chris:openwebui-coding"`2227. Handler calls `api_helpers.clear_workflow_override()` in `finally` block223224### Key Files225226* `api_helpers.py`: Contains `parse_model_field()`, `set_workflow_override()`, `clear_workflow_override()`227* `instance_global_variables.py`: Contains `WORKFLOW_OVERRIDE` global variable228* `workflow_gateway.py`: Checks for workflow override before normal routing229* `config_utils.py`: Contains `get_shared_workflows_folder()`, `get_available_shared_workflows()`,230 `workflow_exists_in_shared_folder()`231232-----233234## 5\. How to Extend235236The architecture makes it easy to add new functionality. The `ApiServer` automatically discovers new handlers, so in237most cases, you only need to add new files without modifying existing ones.238239### **Example 2: Add Support for a New API Type (e.g., Anthropic)**240241This more advanced example shows how to add compatibility for an entirely new API schema, supporting both streaming and242non-streaming.2432441. **Update `ResponseBuilderService`**: Add methods to build both the final response and the streaming chunks for the245 new API type.2462472. **Update `api_helpers` for Streaming**: Wire the new streaming chunk builder into the streaming dispatcher248 in `build_response_json`.2492503. **Create the New Handler File**: Create the file that will define the endpoints and translation logic for the new251 API.252253 * Create `Middleware/api/handlers/impl/anthropic_api_handler.py`.254 * Implement the full logic. This includes generating a `request_id`, parsing the incoming request, setting255 the `API_TYPE`, calling the gateway, and using the correct response builder methods for both streaming and256 non-streaming cases.257258 <!-- end list -->259260 ```python261 # In Middleware/api/handlers/impl/anthropic_api_handler.py262 import uuid263 from flask import request, jsonify, Response, g264 from flask.views import MethodView265 from Middleware.api.app import app266 from Middleware.api.handlers.base.base_api_handler import BaseApiHandler267 from Middleware.api.workflow_gateway import handle_user_prompt268 from Middleware.common import instance_global_variables269 from Middleware.services.response_builder_service import ResponseBuilderService270271 response_builder = ResponseBuilderService()272273 class AnthropicChatAPI(MethodView):274 @staticmethod275 def post() -> Response:276 # 1. Generate request ID for cancellation tracking277 request_id = str(uuid.uuid4())278 g.current_request_id = request_id279280 # 2. Set the API type so downstream services know the format281 instance_global_variables.API_TYPE = "anthropicchat"282 data = request.get_json()283 284 # 3. Transform the Anthropic request into our internal format285 messages = [{"role": m["role"], "content": m["content"][0]["text"]} for m in data.get("messages", [])]286 287 stream = data.get("stream", True)288289 # 4. Handoff to the workflow engine290 llm_response = handle_user_prompt(request_id, messages, stream=stream)291292 if stream:293 return Response(llm_response, mimetype='text/event-stream')294 else:295 response_payload = response_builder.build_anthropic_response(llm_response)296 return jsonify(response_payload)297298 class AnthropicApiHandler(BaseApiHandler):299 def register_routes(self, app_instance: app):300 app_instance.add_url_rule(301 '/v1/anthropic/messages', 302 view_func=AnthropicChatAPI.as_view('anthropic_chat_api')303 )304 ```