1---2name: computer-vision-api3description: Computer Vision API skill. Use when working with Computer Vision for models, analyze, generateThumbnail. Covers 9 endpoints.4---5
6# Computer Vision API
7API version: 1.0
8
9## Auth
10ApiKey Ocp-Apim-Subscription-Key in header
11
12## Base URL
13Not specified.
14
15## Setup
161. Set your API key in the appropriate header
172. GET /models -- verify access
183. POST /analyze -- create first analyze
19
20## Endpoints
21
229 endpoints across 8 groups. See references/api-spec.lap for full details.
23
24### models
25| Method | Path | Description |
26|--------|------|-------------|
27| GET | /models | This operation returns the list of domain-specific models that are supported by the Computer Vision API. Currently, the API only supports one domain-specific model: a celebrity recognizer. A successful response will be returned in JSON. If the request failed, the response will contain an error code and a message to help understand what went wrong. |
28| POST | /models/{model}/analyze | This operation recognizes content within an image by applying a domain-specific model. The list of domain-specific models that are supported by the Computer Vision API can be retrieved using the /models GET request. Currently, the API only provides a single domain-specific model: celebrities. Two input methods are supported -- (1) Uploading an image or (2) specifying an image URL. A successful response will be returned in JSON. If the request failed, the response will contain an error code and a message to help understand what went wrong. |
29
30### analyze
31| Method | Path | Description |
32|--------|------|-------------|
33| POST | /analyze | This operation extracts a rich set of visual features based on the image content. Two input methods are supported -- (1) Uploading an image or (2) specifying an image URL. Within your request, there is an optional parameter to allow you to choose which features to return. By default, image categories are returned in the response. |
34
35### generateThumbnail
36| Method | Path | Description |
37|--------|------|-------------|
38| POST | /generateThumbnail | This operation generates a thumbnail image with the user-specified width and height. By default, the service analyzes the image, identifies the region of interest (ROI), and generates smart cropping coordinates based on the ROI. Smart cropping helps when you specify an aspect ratio that differs from that of the input image. A successful response contains the thumbnail image binary. If the request failed, the response contains an error code and a message to help determine what went wrong. |
39
40### ocr
41| Method | Path | Description |
42|--------|------|-------------|
43| POST | /ocr | Optical Character Recognition (OCR) detects printed text in an image and extracts the recognized characters into a machine-usable character stream. Upon success, the OCR results will be returned. Upon failure, the error code together with an error message will be returned. The error code can be one of InvalidImageUrl, InvalidImageFormat, InvalidImageSize, NotSupportedImage, NotSupportedLanguage, or InternalServerError. |
44
45### describe
46| Method | Path | Description |
47|--------|------|-------------|
48| POST | /describe | This operation generates a description of an image in human readable language with complete sentences. The description is based on a collection of content tags, which are also returned by the operation. More than one description can be generated for each image. Descriptions are ordered by their confidence score. All descriptions are in English. Two input methods are supported -- (1) Uploading an image or (2) specifying an image URL.A successful response will be returned in JSON. If the request failed, the response will contain an error code and a message to help understand what went wrong. |
49
50### tag
51| Method | Path | Description |
52|--------|------|-------------|
53| POST | /tag | This operation generates a list of words, or tags, that are relevant to the content of the supplied image. The Computer Vision API can return tags based on objects, living beings, scenery or actions found in images. Unlike categories, tags are not organized according to a hierarchical classification system, but correspond to image content. Tags may contain hints to avoid ambiguity or provide context, for example the tag 'cello' may be accompanied by the hint 'musical instrument'. All tags are in English. |
54
55### recognizeText
56| Method | Path | Description |
57|--------|------|-------------|
58| POST | /recognizeText | Recognize Text operation. When you use the Recognize Text interface, the response contains a field called 'Operation-Location'. The 'Operation-Location' field contains the URL that you must use for your Get Handwritten Text Operation Result operation. |
59
60### textOperations
61| Method | Path | Description |
62|--------|------|-------------|
63| GET | /textOperations/{operationId} | This interface is used for getting text operation result. The URL to this interface should be retrieved from 'Operation-Location' field returned from Recognize Text interface. |
64
65## Common Questions
66
67Match user requests to endpoints in references/api-spec.lap. Key patterns:
68- "List all models?" -> GET /models
69- "Create a analyze?" -> POST /analyze
70- "Create a generateThumbnail?" -> POST /generateThumbnail
71- "Create a ocr?" -> POST /ocr
72- "Create a describe?" -> POST /describe
73- "Create a tag?" -> POST /tag
74- "Create a analyze?" -> POST /models/{model}/analyze
75- "Create a recognizeText?" -> POST /recognizeText
76- "Get textOperation details?" -> GET /textOperations/{operationId}
77- "How to authenticate?" -> See Auth section
78
79## Response Tips
80- Check response schemas in references/api-spec.lap for field details
81- Create/update endpoints typically return the created/updated object
82
83## References
84- Full spec: See references/api-spec.lap for complete endpoint details, parameter tables, and response schemas
85
86> Generated from the official API spec by [LAP](https://lap.sh)