Serve local model endpoints for agent tests with OpenLLM
Launch an OpenAI-compatible OpenLLM server for a chosen open model, point an agent runtime at it, and compare behavior before production use.
Prerequisites
OpenLLM, Python environment, supported open model, required GPU/CPU resources, Hugging Face token for gated models, agent runtime or SDK that can call an OpenAI-compatible endpoint.
Installation
Use the upstream install or setup path that matches your environment:
- pip install openllm # or pip3 install openllm
Requirements and caveats from upstream:
- python
Basic usage or getting-started notes:
OpenLLM allows developers to run any open-source LLMs (Llama 3.3, Qwen2.5, Phi3 and more) or custom models as OpenAI-compatible APIs with a single command. It features a [built-in chat...
Run the following commands to install OpenLLM and explore it interactively.
OpenLLM supports a wide range of state-of-the-art open-source LLMs. You can also add a model repository to run custom models with OpenLLM.
Extracted from upstream docs: https://raw.githubusercontent.com/bentoml/OpenLLM/HEAD/README.md