import Image from '@theme/IdealImage';
Benchmark LLMs
Easily benchmark LLMs for a given question by viewing
- Responses
- Response Cost
- Response Time
Benchmark Output
<Image img={require('../../img/bench_llm.png')} />
Setup:
git clone https://github.com/BerriAI/litellm
cd to litellm/cookbook/benchmark dir
Located here:
https://github.com/BerriAI/litellm/tree/main/cookbook/benchmark
cd litellm/cookbook/benchmark
Install Dependencies
pip install litellm click tqdm tabulate termcolor
Configuration - Set LLM API Keys + LLMs in benchmark.py
In benchmark/benchmark.py select your LLMs, LLM API Key and questions
Supported LLMs: https://docs.litellm.ai/docs/providers
# Define the list of models to benchmark
models = ['gpt-3.5-turbo', 'claude-2']
# Enter LLM API keys
os.environ['OPENAI_API_KEY'] = ""
os.environ['ANTHROPIC_API_KEY'] = ""
# List of questions to benchmark (replace with your questions)
questions = [
"When will BerriAI IPO?",
"When will LiteLLM hit $100M ARR?"
]
Run benchmark.py
python3 benchmark.py
Expected Output
Running question: When will BerriAI IPO? for model: claude-2: 100%|████████████████████████████████████████████████████████████████████████████████████| 3/3 [00:13<00:00, 4.41s/it]
Benchmark Results for 'When will BerriAI IPO?':
+-----------------+----------------------------------------------------------------------------------+---------------------------+------------+
| Model | Response | Response Time (seconds) | Cost ($) |
+=================+==================================================================================+===========================+============+
| gpt-3.5-turbo | As an AI language model, I cannot provide up-to-date information or predict | 1.55 seconds | $0.000122 |
| | future events. It is best to consult a reliable financial source or contact | | |
| | BerriAI directly for information regarding their IPO plans. | | |
+-----------------+----------------------------------------------------------------------------------+---------------------------+------------+
| togethercompute | I'm not able to provide information about future IPO plans or dates for BerriAI | 8.52 seconds | $0.000531 |
| r/llama-2-70b-c | or any other company. IPO (Initial Public Offering) plans and timelines are | | |
| hat | typically kept private by companies until they are ready to make a public | | |
| | announcement. It's important to note that IPO plans can change and are subject | | |
| | to various factors, such as market conditions, financial performance, and | | |
| | regulatory approvals. Therefore, it's difficult to predict with certainty when | | |
| | BerriAI or any other company will go public. If you're interested in staying | | |
| | up-to-date with BerriAI's latest news and developments, you may want to follow | | |
| | their official social media accounts, subscribe to their newsletter, or visit | | |
| | their website periodically for updates. | | |
+-----------------+----------------------------------------------------------------------------------+---------------------------+------------+
| claude-2 | I do not have any information about when or if BerriAI will have an initial | 3.17 seconds | $0.002084 |
| | public offering (IPO). As an AI assistant created by Anthropic to be helpful, | | |
| | harmless, and honest, I do not have insider knowledge about Anthropic's business | | |
| | plans or strategies. | | |
+-----------------+----------------------------------------------------------------------------------+---------------------------+------------+
Support
🤝 Schedule a 1-on-1 Session: Book a 1-on-1 session with Krrish and Ishaan, the founders, to discuss any issues, provide feedback, or explore how we can improve LiteLLM for you.
1---2name: benchmark-llms3description: In benchmark/benchmark.py select your LLMs, LLM API Key and questions4---5import Image from '@theme/IdealImage';67# Benchmark LLMs8Easily benchmark LLMs for a given question by viewing 9* Responses 10* Response Cost11* Response Time1213### Benchmark Output14<Image img={require('../../img/bench_llm.png')} />1516## Setup:17```18git clone https://github.com/BerriAI/litellm19```20cd to `litellm/cookbook/benchmark` dir2122Located here: 23https://github.com/BerriAI/litellm/tree/main/cookbook/benchmark24```25cd litellm/cookbook/benchmark26```2728### Install Dependencies29```30pip install litellm click tqdm tabulate termcolor31```3233### Configuration - Set LLM API Keys + LLMs in benchmark.py34In `benchmark/benchmark.py` select your LLMs, LLM API Key and questions3536Supported LLMs: https://docs.litellm.ai/docs/providers3738```python39# Define the list of models to benchmark40models = ['gpt-3.5-turbo', 'claude-2']4142# Enter LLM API keys43os.environ['OPENAI_API_KEY'] = ""44os.environ['ANTHROPIC_API_KEY'] = ""4546# List of questions to benchmark (replace with your questions)47questions = [48 "When will BerriAI IPO?",49 "When will LiteLLM hit $100M ARR?"50]5152```5354## Run benchmark.py55```56python3 benchmark.py57```5859## Expected Output60```61Running question: When will BerriAI IPO? for model: claude-2: 100%|████████████████████████████████████████████████████████████████████████████████████| 3/3 [00:13<00:00, 4.41s/it]6263Benchmark Results for 'When will BerriAI IPO?':64+-----------------+----------------------------------------------------------------------------------+---------------------------+------------+65| Model | Response | Response Time (seconds) | Cost ($) |66+=================+==================================================================================+===========================+============+67| gpt-3.5-turbo | As an AI language model, I cannot provide up-to-date information or predict | 1.55 seconds | $0.000122 |68| | future events. It is best to consult a reliable financial source or contact | | |69| | BerriAI directly for information regarding their IPO plans. | | |70+-----------------+----------------------------------------------------------------------------------+---------------------------+------------+71| togethercompute | I'm not able to provide information about future IPO plans or dates for BerriAI | 8.52 seconds | $0.000531 |72| r/llama-2-70b-c | or any other company. IPO (Initial Public Offering) plans and timelines are | | |73| hat | typically kept private by companies until they are ready to make a public | | |74| | announcement. It's important to note that IPO plans can change and are subject | | |75| | to various factors, such as market conditions, financial performance, and | | |76| | regulatory approvals. Therefore, it's difficult to predict with certainty when | | |77| | BerriAI or any other company will go public. If you're interested in staying | | |78| | up-to-date with BerriAI's latest news and developments, you may want to follow | | |79| | their official social media accounts, subscribe to their newsletter, or visit | | |80| | their website periodically for updates. | | |81+-----------------+----------------------------------------------------------------------------------+---------------------------+------------+82| claude-2 | I do not have any information about when or if BerriAI will have an initial | 3.17 seconds | $0.002084 |83| | public offering (IPO). As an AI assistant created by Anthropic to be helpful, | | |84| | harmless, and honest, I do not have insider knowledge about Anthropic's business | | |85| | plans or strategies. | | |86+-----------------+----------------------------------------------------------------------------------+---------------------------+------------+87```88## Support89**🤝 Schedule a 1-on-1 Session:** Book a [1-on-1 session](https://calendly.com/d/cx9p-5yf-2nm/litellm-introductions) with Krrish and Ishaan, the founders, to discuss any issues, provide feedback, or explore how we can improve LiteLLM for you.909192<!-- 93## Pre-requisites:94``` python95!pip install litellm96```9798## Example Use Case 1 - Code Generator99100### Enter your system prompt and questions101```` python102# enter your system prompt if you have one103system_prompt = """104You are a coding assistant helping users using litellm.105litellm is a light package to simplify calling OpenAI, Azure, Cohere, Anthropic, Huggingface API Endpoints106--107Sample Usage:108```109pip install litellm110from litellm import completion111## set ENV variables112os.environ["OPENAI_API_KEY"] = "openai key"113os.environ["COHERE_API_KEY"] = "cohere key"114messages = [{ "content": "Hello, how are you?","role": "user"}]115# openai call116response = completion(model="gpt-3.5-turbo", messages=messages)117# cohere call118response = completion("command-nightly", messages)119```120121"""122123124# questions/logs you want to run the LLM on125questions = [126 "what is litellm?",127 "why should I use LiteLLM",128 "does litellm support Anthropic LLMs",129 "write code to make a litellm completion call",130]131````132133### Running questions134135### Select from 100+ LLMs here: <https://docs.litellm.ai/docs/providers> {#select-from-100-llms-here-httpsdocslitellmaidocsproviders}136137``` python138import litellm139from litellm import completion, completion_cost140import os141import time142143# optional use litellm dashboard to view logs144# litellm.use_client = True145# litellm.token = "ishaan_2@berri.ai" # set your email146147148# set API keys149os.environ['TOGETHERAI_API_KEY'] = ""150os.environ['OPENAI_API_KEY'] = ""151os.environ['ANTHROPIC_API_KEY'] = ""152153154# select LLMs to benchmark155# using https://api.together.xyz/playground for llama2156# try any supported LLM here: https://docs.litellm.ai/docs/providers157158models = ['togethercomputer/llama-2-70b-chat', 'gpt-3.5-turbo', 'claude-instant-1.2']159data = []160161for question in questions: # group by question162 for model in models:163 print(f"running question: {question} for model: {model}")164 start_time = time.time()165 # show response, response time, cost for each question166 response = completion(167 model=model,168 max_tokens=500,169 messages = [170 {171 "role": "system", "content": system_prompt172 },173 {174 "role": "user", "content": question175 }176 ],177 )178 end = time.time()179 total_time = end-start_time # response time180 # print(response)181 cost = completion_cost(response) # cost for completion182 raw_response = response['choices'][0]['message']['content'] # response string183184185 # add log to pandas df186 data.append(187 {188 'Model': model,189 'Question': question,190 'Response': raw_response,191 'ResponseTime': total_time,192 'Cost': cost193 })194```195196### View Benchmarks for LLMs197``` python198from IPython.display import display199from IPython.core.interactiveshell import InteractiveShell200InteractiveShell.ast_node_interactivity = "all"201from IPython.display import HTML202import pandas as pd203204df = pd.DataFrame(data)205grouped_by_question = df.groupby('Question')206207for question, group_data in grouped_by_question:208 print(f"Question: {question}")209 HTML(group_data.to_html())210```211212<table border="1" class="dataframe">213 <thead>214 <tr>215 <th></th>216 <th>Model</th>217 <th>Question</th>218 <th>Response</th>219 <th>ResponseTime</th>220 <th>Cost</th>221 </tr>222 </thead>223 <tbody>224 <tr>225 <th>0</th>226 <td>togethercomputer/llama-2-70b-chat</td>227 <td>what is litellm?</td>228 <td>Litellm is a lightweight Python package that simplifies calling various AI API endpoints, including OpenAI, Azure, Cohere, Anthropic, and Hugging Face. It provides a convenient interface for making requests to these APIs, allowing developers to easily integrate them into their applications. With Litellm, developers can quickly and easily interact with multiple AI models and services, without having to handle the details of authentication, API calls, and response parsing. This makes it easier to build and deploy AI-powered applications, and can help developers save time and effort.</td>229 <td>13.479644</td>230 <td>0.000870</td>231 </tr>232 <tr>233 <th>1</th>234 <td>gpt-3.5-turbo</td>235 <td>what is litellm?</td>236 <td>litellm is a light package that provides a simplified interface for making API calls to various language models and APIs. It abstracts away the complexities of handling network requests, authentication, and response parsing, making it easier for developers to integrate powerful language models into their applications.\n\nWith litellm, you can quickly make API calls to models like OpenAI's GPT-3.5 Turbo, Azure's Text Analytics, Cohere's Command API, Anthropic's API, and Huggingface's models. It also supports additional functionality like conversational AI, summarization, translation, and more.\n\nBy using litellm, you can focus on your application logic without getting tangled in the details of API integration, allowing you to quickly build intelligent and conversational applications.</td>237 <td>8.324332</td>238 <td>0.000566</td>239 </tr>240 <tr>241 <th>2</th>242 <td>claude-instant-1.2</td>243 <td>what is litellm?</td>244 <td>litellm is a Python library that simplifies calling various AI API endpoints like OpenAI, Azure, Cohere, Anthropic, and Huggingface. \n\nSome key things to know about litellm:\n\n- It provides a consistent interface for completing prompts and generating responses from different AI models through a single method called completion().\n\n- You specify the API (e.g. OpenAI, Cohere etc.) and model either by name or by setting environment variables before making the completion call.\n\n- This avoids having to use different SDKs or APIs for each provider and standardizes the call structure. \n\n- It handles things like setting headers, encoding inputs, parsing responses so the user doesn't have to deal with those details.\n\n- The goal is to make it easy to try different AI APIs and models without having to change code or learn different interfaces.\n\n- It's lightweight with no other dependencies required besides what's needed for each API (e.g. openai, azure SDKs etc.).\n\nSo in summary, litellm is a small library that provides a common way to interact with multiple conversational AI APIs through a single Python method, avoiding the need to directly use each provider's specific SDK.</td>245 <td>10.316488</td>246 <td>0.001603</td>247 </tr>248 </tbody>249</table>250251## Example Use Case 2 - Rewrite user input concisely252253``` python254# enter your system prompt if you have one255system_prompt = """256For a given user input, rewrite the input to make be more concise.257"""258259# user input for re-writing questions260questions = [261 "LiteLLM is a lightweight Python package that simplifies the process of making API calls to various language models. Here are some reasons why you should use LiteLLM:nn1. **Simplified API Calls**: LiteLLM abstracts away the complexity of making API calls to different language models. It provides a unified interface for invoking models from OpenAI, Azure, Cohere, Anthropic, Huggingface, and more.nn2. **Easy Integration**: LiteLLM seamlessly integrates with your existing codebase. You can import the package and start making API calls with just a few lines of code.nn3. **Flexibility**: LiteLLM supports a variety of language models, including GPT-3, GPT-Neo, chatGPT, and more. You can choose the model that suits your requirements and easily switch between them.nn4. **Convenience**: LiteLLM handles the authentication and connection details for you. You just need to set the relevant environment variables, and the package takes care of the rest.nn5. **Quick Prototyping**: LiteLLM is ideal for rapid prototyping and experimentation. With its simple API, you can quickly generate text, chat with models, and build interactive applications.nn6. **Community Support**: LiteLLM is actively maintained and supported by a community of developers. You can find help, share ideas, and collaborate with others to enhance your projects.nnOverall, LiteLLM simplifies the process of making API calls to language models, saving you time and effort while providing flexibility and convenience",262 "Hi everyone! I'm [your name] and I'm currently working on [your project/role involving LLMs]. I came across LiteLLM and was really excited by how it simplifies working with different LLM providers. I'm hoping to use LiteLLM to [build an app/simplify my code/test different models etc]. Before finding LiteLLM, I was struggling with [describe any issues you faced working with multiple LLMs]. With LiteLLM's unified API and automatic translation between providers, I think it will really help me to [goals you have for using LiteLLM]. Looking forward to being part of this community and learning more about how I can build impactful applications powered by LLMs!Let me know if you would like me to modify or expand on any part of this suggested intro. I'm happy to provide any clarification or additional details you need!",263 "Traceloop is a platform for monitoring and debugging the quality of your LLM outputs. It provides you with a way to track the performance of your LLM application; rollout changes with confidence; and debug issues in production. It is based on OpenTelemetry, so it can provide full visibility to your LLM requests, as well vector DB usage, and other infra in your stack."264]265```266267### Run Questions268269``` python270import litellm271from litellm import completion, completion_cost272import os273import time274275# optional use litellm dashboard to view logs276# litellm.use_client = True277# litellm.token = "ishaan_2@berri.ai" # set your email278279os.environ['TOGETHERAI_API_KEY'] = ""280os.environ['OPENAI_API_KEY'] = ""281os.environ['ANTHROPIC_API_KEY'] = ""282283models = ['togethercomputer/llama-2-70b-chat', 'gpt-3.5-turbo', 'claude-instant-1.2'] # enter llms to benchmark284data_2 = []285286for question in questions: # group by question287 for model in models:288 print(f"running question: {question} for model: {model}")289 start_time = time.time()290 # show response, response time, cost for each question291 response = completion(292 model=model,293 max_tokens=500,294 messages = [295 {296 "role": "system", "content": system_prompt297 },298 {299 "role": "user", "content": "User input:" + question300 }301 ],302 )303 end = time.time()304 total_time = end-start_time # response time305 # print(response)306 cost = completion_cost(response) # cost for completion307 raw_response = response['choices'][0]['message']['content'] # response string308 #print(raw_response, total_time, cost)309310 # add to pandas df311 data_2.append(312 {313 'Model': model,314 'Question': question,315 'Response': raw_response,316 'ResponseTime': total_time,317 'Cost': cost318 })319320321```322### View Logs - Group by Question323``` python324from IPython.display import display325from IPython.core.interactiveshell import InteractiveShell326InteractiveShell.ast_node_interactivity = "all"327from IPython.display import HTML328import pandas as pd329330df = pd.DataFrame(data_2)331grouped_by_question = df.groupby('Question')332333for question, group_data in grouped_by_question:334 print(f"Question: {question}")335 HTML(group_data.to_html())336```337338#### User Question339 Question: Hi everyone! I'm [your name] and I'm currently working on [your project/role involving LLMs]. I came across LiteLLM and was really excited by how it simplifies working with different LLM providers. I'm hoping to use LiteLLM to [build an app/simplify my code/test different models etc]. Before finding LiteLLM, I was struggling with [describe any issues you faced working with multiple LLMs]. With LiteLLM's unified API and automatic translation between providers, I think it will really help me to [goals you have for using LiteLLM]. Looking forward to being part of this community and learning more about how I can build impactful applications powered by LLMs!Let me know if you would like me to modify or expand on any part of this suggested intro. I'm happy to provide any clarification or additional details you need!340#### Logs341<table border="1" class="dataframe">342 <thead>343 <tr>344 <th></th>345 <th>Model</th>346 <th>Response</th>347 <th>ResponseTime</th>348 <th>Cost</th>349 </tr>350 </thead>351 <tbody>352 <tr>353 <th>3</th>354 <td>togethercomputer/llama-2-70b-chat</td>355 <td>nHere's a more concise version of the user input:nn"Hi everyone! I'm [your name] and I'm working on [your project/role involving LLMs]. I recently discovered LiteLLM and I'm excited to use it to [build an app/simplify my code/test different models etc]. Before LiteLLM, I struggled with [describe any issues you faced working with multiple LLMs]. I'm looking forward to using LiteLLM's unified API and automatic translation to achieve my goals. I'm eager to learn more about building impactful applications powered by LLMs and to be part of this community. Let me know if you have any questions or need further clarification."nnIn this revised version, we've kept the essential information and removed some of the extraneous language. We've also rephrased some of the sentences to make them more concise and easier to read.</td>356 <td>18.300620</td>357 <td>0.001200</td>358 </tr>359 <tr>360 <th>4</th>361 <td>gpt-3.5-turbo</td>362 <td>User input: Hi, I'm [your name] and I'm excited about using LiteLLM to simplify working with different LLM providers. Before finding LiteLLM, I faced challenges working with multiple LLMs. With LiteLLM's unified API and automatic translation, I believe it will help me achieve my goals of [state your goals]. I look forward to being part of this community and learning how to build impactful applications with LLMs. Let me know if you need any further clarification or details.</td>363 <td>7.385472</td>364 <td>0.000525</td>365 </tr>366 <tr>367 <th>5</th>368 <td>claude-instant-1.2</td>369 <td>Here is a more concise rewrite of the user input:nnHi everyone, I'm [your name]. I'm currently [your project/role] and came across LiteLLM, which simplifies working with different LLMs through its unified API. I hope to [build an app/simplify code/test models] with LiteLLM since I previously struggled with [issues]. LiteLLM's automatic translation between providers will help me [goals] and build impactful LLM applications. Looking forward to learning more as part of this community. Let me know if you need any clarification on my plans to use LiteLLM.</td>370 <td>8.628217</td>371 <td>0.001022</td>372 </tr>373 </tbody>374</table> -->