Endpoint Configuration Files
The Endpoint configuration files define a connection to a specific Large Language Model (LLM) backend. Each JSON file in
the Public/Configs/Endpoints/ directory tells the WilmerAI middleware the server's address, its API format, and a set
of rules for injecting text into prompts and cleaning the model's raw responses.
Field Definitions
Each Endpoint JSON file contains a single object with the following key-value pairs, organized by function.
Core Connection Details
These fields are essential for establishing a connection to the LLM server.
endpoint
- Description: The full URL of the LLM API server, such as an OpenAI-compatible service.
- Data Type:
string
- Required: Yes
- Example:
"http://localhost:11434"
apiTypeConfigFileName
- Description: The name of the JSON file (without the
.json extension) from the Public/Configs/ApiTypes/
directory. This file defines the specific API schema (e.g., Ollama, KoboldCpp, OpenAI) that WilmerAI should use for
this endpoint.
- Data Type:
string
- Required: Yes
- Example:
"Ollama"
maxContextTokenSize
- Description: The maximum context size in tokens that the model can handle. This value is used by the corresponding
ApiTypes configuration to set the model's truncation length property.
- Data Type:
integer
- Required: Yes
- Example:
8192
Model Identification
These fields control how the model is identified in the API request.
modelNameForDisplayOnly
- Description: A human-readable name for this endpoint configuration. This field is for your reference only and is *
not used by the application logic*.
- Data Type:
string
- Required: No
- Example:
"Local Llama3 for Summaries"
modelNameToSendToAPI
- Description: The specific model identifier to be included in the API request payload. This is necessary for
multi-model servers like Ollama or OpenAI.
- Data Type:
string
- Required: Yes
- Example:
"llama3:8b-instruct-q5_K_M"
dontIncludeModel
- Description: If
true, the modelNameToSendToAPI field and its corresponding key will be omitted from the API
request payload. This is useful for single-model servers that do not accept a model parameter.
- Data Type:
boolean
- Required: Yes
- Example:
true
Prompt & Content Injection
These settings allow for adding text to different parts of the prompt before it is sent to the LLM.
addTextToStartOfSystem
- Description: A flag that, when
true, prepends the content of textToAddToStartOfSystem to the system prompt.
- Data Type:
boolean
- Required: Yes
- Example:
true
textToAddToStartOfSystem
- Description: The string content to prepend to the system prompt when
addTextToStartOfSystem is enabled.
- Data Type:
string
- Required: Yes
- Example:
"/no_think "
addTextToStartOfPrompt
- Description: A flag that, when
true, prepends the content of textToAddToStartOfPrompt to the final user
prompt. For Chat Completion APIs, this modifies the content of the last message with role: "user".
- Data Type:
boolean
- Required: Yes
- Example:
false
textToAddToStartOfPrompt
- Description: The string content to prepend to the user prompt when
addTextToStartOfPrompt is enabled.
- Data Type:
string
- Required: Yes
- Example:
""
addTextToStartOfCompletion
- Description: A flag that, when
true, appends the content of textToAddToStartOfCompletion to the end of the
context to "seed" the AI's response, forcing it to begin its generation with the specified text.
- Data Type:
boolean
- Required: Yes
- Example:
true
textToAddToStartOfCompletion
- Description: The string content used to seed the AI's response when
addTextToStartOfCompletion is enabled.
- Data Type:
string
- Required: Yes
- Example:
"<think>First, I will analyze the user's request.</think>"
ensureTextAddedToAssistantWhenChatCompletion
- Description: Modifies the behavior of
addTextToStartOfCompletion for Chat Completion APIs. If true and the
last message is not from the assistant, a new role: "assistant" message is added containing the seed text. If
false, the text is appended to the content of the last message, regardless of its role.
- Data Type:
boolean
- Required: Yes
- Example:
true
Response Cleaning & Filtering
These settings process the raw text from the LLM after it is received. Operations are applied in this order: 1) Thinking
Block Removal, 2) Custom Prefix Removal, 3) Whitespace Trimming.
removeThinking
- Description: The master switch for the thinking block removal feature. If
true, the system will attempt to find
and remove text between startThinkTag and endThinkTag.
- Data Type:
boolean
- Required: Yes
- Example:
true
startThinkTag
- Description: The opening tag that marks the beginning of a "thinking" block to be removed from the response.
Case-insensitive.
- Data Type:
string
- Required: Yes
- Example:
"<think>"
endThinkTag
- Description: The closing tag that marks the end of a "thinking" block. Case-insensitive.
- Data Type:
string
- Required: Yes
- Example:
"</think>"
openingTagGracePeriod
- Description: The number of characters at the start of the response to scan for a
startThinkTag. If the tag is
not found within this window, removal is skipped for that response.
- Data Type:
integer
- Required: Yes
- Example:
100
expectOnlyClosingThinkTag
- Description: A special mode for models that may omit the opening tag. If
true, the system buffers and discards
all text until it finds the endThinkTag, then begins streaming the subsequent text.
- Data Type:
boolean
- Required: Yes
- Example:
false
removeCustomTextFromResponseStartEndpointWide
- Description: If
true, enables the removal of one of the strings defined in
responseStartTextToRemoveEndpointWide from the beginning of the LLM's response.
- Data Type:
boolean
- Required: Yes
- Example:
false
responseStartTextToRemoveEndpointWide
- Description: A list of strings to remove from the beginning of the response. The system removes the first string
in the list that it finds a match for.
- Data Type:
list of strings
- Required: Yes
- Example:
["Assistant:", "Okay, here is the information you requested:"]
trimBeginningAndEndLineBreaks
- Description: If
true, any leading or trailing whitespace (spaces, newlines, tabs) is removed from the final
response.
- -Data Type:
boolean
- Required: Yes
- Example:
true
Example Endpoint File
{
"modelNameForDisplayOnly": "Small model for all tasks",
"endpoint": "http://12.0.0.1:5000",
"apiTypeConfigFileName": "KoboldCpp",
"maxContextTokenSize": 8192,
"modelNameToSendToAPI": "",
"trimBeginningAndEndLineBreaks": true,
"dontIncludeModel": false,
"removeThinking": true,
"startThinkTag": "<think>",
"endThinkTag": "</think>",
"openingTagGracePeriod": 100,
"expectOnlyClosingThinkTag": false,
"addTextToStartOfSystem": true,
"textToAddToStartOfSystem": "/no_think ",
"addTextToStartOfPrompt": false,
"textToAddToStartOfPrompt": "",
"addTextToStartOfCompletion": false,
"textToAddToStartOfCompletion": "",
"ensureTextAddedToAssistantWhenChatCompletion": false,
"removeCustomTextFromResponseStartEndpointWide": false,
"responseStartTextToRemoveEndpointWide": []
}
1---2name: endpoint-configuration-files3description: The Endpoint configuration files define a connection to a specific Large Language Model (LLM) backend.4---5### **`Endpoint` Configuration Files**67The Endpoint configuration files define a connection to a specific Large Language Model (LLM) backend. Each JSON file in8the `Public/Configs/Endpoints/` directory tells the WilmerAI middleware the server's address, its API format, and a set9of rules for injecting text into prompts and cleaning the model's raw responses.1011-----1213#### **Field Definitions**1415Each Endpoint JSON file contains a single object with the following key-value pairs, organized by function.1617-----1819#### **Core Connection Details**2021These fields are essential for establishing a connection to the LLM server.2223##### `endpoint`2425* **Description**: The full URL of the LLM API server, such as an OpenAI-compatible service.26* **Data Type**: `string`27* **Required**: Yes28* **Example**: `"http://localhost:11434"`2930##### `apiTypeConfigFileName`3132* **Description**: The name of the JSON file (without the `.json` extension) from the `Public/Configs/ApiTypes/`33 directory. This file defines the specific API schema (e.g., Ollama, KoboldCpp, OpenAI) that WilmerAI should use for34 this endpoint.35* **Data Type**: `string`36* **Required**: Yes37* **Example**: `"Ollama"`3839##### `maxContextTokenSize`4041* **Description**: The maximum context size in tokens that the model can handle. This value is used by the corresponding42 `ApiTypes` configuration to set the model's truncation length property.43* **Data Type**: `integer`44* **Required**: Yes45* **Example**: `8192`4647-----4849#### **Model Identification**5051These fields control how the model is identified in the API request.5253##### `modelNameForDisplayOnly`5455* **Description**: A human-readable name for this endpoint configuration. This field is for your reference only and is *56 *not used by the application logic**.57* **Data Type**: `string`58* **Required**: No59* **Example**: `"Local Llama3 for Summaries"`6061##### `modelNameToSendToAPI`6263* **Description**: The specific model identifier to be included in the API request payload. This is necessary for64 multi-model servers like Ollama or OpenAI.65* **Data Type**: `string`66* **Required**: Yes67* **Example**: `"llama3:8b-instruct-q5_K_M"`6869##### `dontIncludeModel`7071* **Description**: If `true`, the `modelNameToSendToAPI` field and its corresponding key will be omitted from the API72 request payload. This is useful for single-model servers that do not accept a model parameter.73* **Data Type**: `boolean`74* **Required**: Yes75* **Example**: `true`7677-----7879#### **Prompt & Content Injection**8081These settings allow for adding text to different parts of the prompt before it is sent to the LLM.8283##### `addTextToStartOfSystem`8485* **Description**: A flag that, when `true`, prepends the content of `textToAddToStartOfSystem` to the system prompt.86* **Data Type**: `boolean`87* **Required**: Yes88* **Example**: `true`8990##### `textToAddToStartOfSystem`9192* **Description**: The string content to prepend to the system prompt when `addTextToStartOfSystem` is enabled.93* **Data Type**: `string`94* **Required**: Yes95* **Example**: `"/no_think "`9697##### `addTextToStartOfPrompt`9899* **Description**: A flag that, when `true`, prepends the content of `textToAddToStartOfPrompt` to the final user100 prompt. For Chat Completion APIs, this modifies the content of the last message with `role: "user"`.101* **Data Type**: `boolean`102* **Required**: Yes103* **Example**: `false`104105##### `textToAddToStartOfPrompt`106107* **Description**: The string content to prepend to the user prompt when `addTextToStartOfPrompt` is enabled.108* **Data Type**: `string`109* **Required**: Yes110* **Example**: `""`111112##### `addTextToStartOfCompletion`113114* **Description**: A flag that, when `true`, appends the content of `textToAddToStartOfCompletion` to the end of the115 context to "seed" the AI's response, forcing it to begin its generation with the specified text.116* **Data Type**: `boolean`117* **Required**: Yes118* **Example**: `true`119120##### `textToAddToStartOfCompletion`121122* **Description**: The string content used to seed the AI's response when `addTextToStartOfCompletion` is enabled.123* **Data Type**: `string`124* **Required**: Yes125* **Example**: `"<think>First, I will analyze the user's request.</think>"`126127##### `ensureTextAddedToAssistantWhenChatCompletion`128129* **Description**: Modifies the behavior of `addTextToStartOfCompletion` for Chat Completion APIs. If `true` and the130 last message is not from the assistant, a new `role: "assistant"` message is added containing the seed text. If131 `false`, the text is appended to the content of the last message, regardless of its role.132* **Data Type**: `boolean`133* **Required**: Yes134* **Example**: `true`135136-----137138#### **Response Cleaning & Filtering**139140These settings process the raw text from the LLM after it is received. Operations are applied in this order: 1) Thinking141Block Removal, 2) Custom Prefix Removal, 3) Whitespace Trimming.142143##### `removeThinking`144145* **Description**: The master switch for the thinking block removal feature. If `true`, the system will attempt to find146 and remove text between `startThinkTag` and `endThinkTag`.147* **Data Type**: `boolean`148* **Required**: Yes149* **Example**: `true`150151##### `startThinkTag`152153* **Description**: The opening tag that marks the beginning of a "thinking" block to be removed from the response.154 Case-insensitive.155* **Data Type**: `string`156* **Required**: Yes157* **Example**: `"<think>"`158159##### `endThinkTag`160161* **Description**: The closing tag that marks the end of a "thinking" block. Case-insensitive.162* **Data Type**: `string`163* **Required**: Yes164* **Example**: `"</think>"`165166##### `openingTagGracePeriod`167168* **Description**: The number of characters at the start of the response to scan for a `startThinkTag`. If the tag is169 not found within this window, removal is skipped for that response.170* **Data Type**: `integer`171* **Required**: Yes172* **Example**: `100`173174##### `expectOnlyClosingThinkTag`175176* **Description**: A special mode for models that may omit the opening tag. If `true`, the system buffers and discards177 all text until it finds the `endThinkTag`, then begins streaming the subsequent text.178* **Data Type**: `boolean`179* **Required**: Yes180* **Example**: `false`181182##### `removeCustomTextFromResponseStartEndpointWide`183184* **Description**: If `true`, enables the removal of one of the strings defined in185 `responseStartTextToRemoveEndpointWide` from the beginning of the LLM's response.186* **Data Type**: `boolean`187* **Required**: Yes188* **Example**: `false`189190##### `responseStartTextToRemoveEndpointWide`191192* **Description**: A list of strings to remove from the beginning of the response. The system removes the first string193 in the list that it finds a match for.194* **Data Type**: `list of strings`195* **Required**: Yes196* **Example**: `["Assistant:", "Okay, here is the information you requested:"]`197198##### `trimBeginningAndEndLineBreaks`199200* **Description**: If `true`, any leading or trailing whitespace (spaces, newlines, tabs) is removed from the final201 response.202* \-**Data Type**: `boolean`203* **Required**: Yes204* **Example**: `true`205206-----207208#### **Example Endpoint File**209210```json211{212 "modelNameForDisplayOnly": "Small model for all tasks",213 "endpoint": "http://12.0.0.1:5000",214 "apiTypeConfigFileName": "KoboldCpp",215 "maxContextTokenSize": 8192,216 "modelNameToSendToAPI": "",217 "trimBeginningAndEndLineBreaks": true,218 "dontIncludeModel": false,219 "removeThinking": true,220 "startThinkTag": "<think>",221 "endThinkTag": "</think>",222 "openingTagGracePeriod": 100,223 "expectOnlyClosingThinkTag": false,224 "addTextToStartOfSystem": true,225 "textToAddToStartOfSystem": "/no_think ",226 "addTextToStartOfPrompt": false,227 "textToAddToStartOfPrompt": "",228 "addTextToStartOfCompletion": false,229 "textToAddToStartOfCompletion": "",230 "ensureTextAddedToAssistantWhenChatCompletion": false,231 "removeCustomTextFromResponseStartEndpointWide": false,232 "responseStartTextToRemoveEndpointWide": []233}234```