# Databricks Isv Connector Structure

> How to structure a Databricks connector (REST or Python SDK): config, connect, operations, validation. Use when designing or building a new connector.

- Skill: `databricks-solutions/databricks-isv-connector-structure` (Agent Skill)
- Install (CLI): `npx skillmds@latest add databricks-solutions/databricks-isv-connector-structure`
- Raw SKILL.md: https://api.skillmd.com/api/skills/databricks-solutions/databricks-isv-connector-structure/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Integrations & APIs
- Author: databricks-solutions (https://skillmd.com/u/databricks-solutions)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/databricks-solutions/databricks-isv-connector-structure

---


<!-- skill-version: 1.0.0 -->

# Connector Structure (ISV)

Use this skill when **designing or building** a Databricks connector (REST-based or Python SDK-based). It defines a consistent shape so auth, telemetry, and operations stay correct.

## Recommended auth support

**Support all three auth types:** PAT, OAuth M2M, and OAuth U2M.

| Auth type | Use case | Recommended |
|-----------|----------|-------------|
| **PAT** | Testing, development, simple automation | Yes – optional for production but required for many workflows. |
| **OAuth M2M** | Service-to-service, production backends, automated jobs | Yes – recommended for production. |
| **OAuth U2M** | Interactive user sign-in, desktop/UI apps, user context | Yes – app-implemented browser flow or token pass-through. |

Design your config and `connect()` so the user can choose one of these per connection. Do not mix them in a single connection (one connection = one auth type).

## User-Agent (required, connector-level)

**User-Agent is required and is coded at the connector level.** The **end user does not** provide or configure User-Agent. The connector developer sets it in code (e.g. a constant or build-time value). Format: `<isv-name>_<product-name>/<product-version>` (e.g. `AcmePartner_DataConnector/2.1.0`). The connector sends this header on every API/driver call so usage is attributed correctly in Databricks (audit, query history). Do not expose User-Agent (or product/product_version) as a user-configurable connection parameter.

## Reference

- **Adding a Databricks connector to an existing project (no Databricks yet):** [skills/adding-databricks-connector/SKILL.md](../adding-databricks-connector/SKILL.md) – stack choice, where to integrate, minimal steps.
- **Auth patterns:** [skills/rest-api/authentication.md](../rest-api/authentication.md), [skills/python-sdk/authentication.md](../python-sdk/authentication.md), [skills/python-sql-connector/authentication.md](../python-sql-connector/authentication.md), [skills/python-sqlalchemy/authentication.md](../python-sqlalchemy/authentication.md) (SQLAlchemy dialect), [skills/databricks-connect/authentication.md](../databricks-connect/authentication.md) (Databricks Connect), [skills/java-jdbc/authentication.md](../java-jdbc/authentication.md) (Java JDBC OSS: PAT, M2M, U2M; UserAgentEntry).
- **Environment isolation:** [skills/connector-testing/env-isolation.md](../connector-testing/env-isolation.md) — **read before testing.** Covers the `DATABRICKS_AUTH_TYPE` trap (breaks Python SDK, Go SDK, Java SDK), `~/.databrickscfg` DEFAULT profile leakage, and the `env -i` isolation pattern for tests.
- **Checklist:** [skills/connector-testing/integration-checklist.md](../connector-testing/integration-checklist.md).

---

## 1. Config / connection parameters

Use a **single config object** (or connection params) that includes:

| Parameter | Required | Description |
|-----------|----------|-------------|
| `host` | Yes | Workspace URL (e.g. `https://myworkspace.cloud.databricks.com`). **End user provides.** |
| `auth_type` | Yes | One of: `pat` \| `oauth_m2m` \| `oauth_u2m` \| `token` (pass-through). **Recommend supporting all three:** PAT, M2M, U2M. |
| Credentials by auth_type | Yes | For **PAT:** `token`. For **M2M:** `client_id`, `client_secret`. For **U2M:** `access_token` (or run browser flow once and pass token). **End user provides** (except when connector runs browser flow). |
| `warehouse_id` | If using SQL | Required for Statement Execution API. **End user provides** – or the connector can **extract it from `http_path`** (e.g. path `/sql/1.0/warehouses/abc123` → warehouse_id `abc123`). |
| `http_path` | If using SQL connector | SQL Warehouse HTTP path (e.g. `/sql/1.0/warehouses/abc123`). **End user provides.** The connector can derive `warehouse_id` from this path if needed for the Statement Execution API. |
| `redirect_uri` | Optional (U2M localhost) | Default e.g. `http://localhost:8080/callback`. **End user may override.** |
| **Databricks Connect only** | | |
| `serverless_compute_id` | For Connect serverless | Set to `"auto"` (recommended when supported). **End user may enable.** |
| `cluster_id` or `classic_compute_http_path` | For Connect classic | Cluster ID or HTTP path (e.g. `sql/protocolv1/o/<workspace_id>/<cluster_id>`). **End user provides.** One of serverless or classic per connection. |

**User-Agent:** Not a user parameter. The connector sets it in code (connector-level); see "User-Agent (required, connector-level)" above.

**Rule:** One connection = one auth type. Do not accept or mix multiple auth methods in a single config (e.g. do not set both PAT and M2M on the same connection).

---

## User input by auth type

All options the **end user** must (or may) enter for each auth type. **User-Agent is not a user input** – it is coded at the connector level (see "User-Agent (required, connector-level)" above). If the connector runs SQL or Statement Execution, the user may need to provide `warehouse_id` and/or `http_path`. **Note:** The connector can extract `warehouse_id` from `http_path` (e.g. `/sql/1.0/warehouses/abc123` → `abc123`), so the user need not provide both if they already supply `http_path`.

### PAT

| Option | Required | Description |
|--------|----------|-------------|
| `host` | Yes | Databricks workspace URL (e.g. `https://myworkspace.cloud.databricks.com`). |
| `token` | Yes | Personal access token (starts with `dapi...`). User creates in workspace: Settings → Developer → Access tokens. |
| `warehouse_id` | If using SQL / Statement Execution | SQL Warehouse ID. Can be extracted from `http_path` if user provides that instead (e.g. `/sql/1.0/warehouses/abc123` → `abc123`). |
| `http_path` | If using SQL connector | SQL Warehouse HTTP path (e.g. `/sql/1.0/warehouses/abc123`). Connector can derive `warehouse_id` from this. |

### OAuth M2M (client credentials)

| Option | Required | Description |
|--------|----------|-------------|
| `host` | Yes | Databricks workspace URL. |
| `client_id` | Yes | Service principal application (client) ID (UUID). Created in workspace: Settings → Identity and access → Service principals. |
| `client_secret` | Yes | OAuth client secret for that service principal (generate in Service principals UI). |
| `warehouse_id` | If using SQL / Statement Execution | SQL Warehouse ID. Can be extracted from `http_path` if provided. |
| `http_path` | If using SQL connector | SQL Warehouse HTTP path. Connector can derive `warehouse_id` from this. |

### OAuth U2M (user sign-in)

U2M can be implemented in three ways. Collect only the options for the flow you support.

**Option A – External browser (no custom OAuth app)**  
User signs in via browser; connector uses Databricks built-in OAuth app.

| Option | Required | Description |
|--------|----------|-------------|
| `host` | Yes | Databricks workspace URL. |
| `warehouse_id` | If using SQL / Statement Execution | SQL Warehouse ID. Can be extracted from `http_path` if provided. |
| `http_path` | If using SQL connector | SQL Warehouse HTTP path. Connector can derive `warehouse_id` from this. |

*Do not* ask for or use `client_id` / `client_secret` for this flow (M2M service principal client_id is not valid for browser OAuth).

**Option B – Localhost (custom OAuth app)**  
Admin creates an OAuth app in the account with a localhost redirect URI; user signs in and is redirected back to the connector.

| Option | Required | Description |
|--------|----------|-------------|
| `host` | Yes | Databricks workspace URL. |
| `client_id` | Yes | **OAuth application** client ID from the custom app (Settings → Developer → App connections). Not the M2M service principal client_id. |
| `redirect_uri` | Optional | Redirect URI registered for the app. Default: `http://localhost:8080/callback`. |
| `client_secret` | If app is confidential | OAuth app secret, if the app has one. |
| `warehouse_id` | If using SQL / Statement Execution | SQL Warehouse ID. Can be extracted from `http_path` if provided. |
| `http_path` | If using SQL connector | SQL Warehouse HTTP path. Connector can derive `warehouse_id` from this. |

**Option C – Token pass-through (pre-obtained token)**  
User provides an access token already obtained (e.g. from a hosted callback or refresh).

| Option | Required | Description |
|--------|----------|-------------|
| `host` | Yes | Databricks workspace URL. |
| `access_token` | Yes | OAuth access token (from browser flow or refresh). |
| `warehouse_id` | If using SQL / Statement Execution | SQL Warehouse ID. Can be extracted from `http_path` if provided. |
| `http_path` | If using SQL connector | SQL Warehouse HTTP path. Connector can derive `warehouse_id` from this. |

---

## 2. Single entry point: connect(config)

Expose one way to create a connected client:

- **REST connector:** `connect(config) -> Client`, where the client holds the resolved token (or M2M token fetcher), host, and User-Agent. All subsequent API calls use this client so headers are consistent.
- **Python SDK connector:** `connect(config) -> WorkspaceClient` (or `DatabricksSession` for Connect). Build `Config` once from `config` with the correct `auth_type` (e.g. `auth_type="oauth-m2m"` for M2M). Set User-Agent (product/product_version) from connector-level constants, not from user config; return `WorkspaceClient(config=...)`.

Inside `connect()`:

1. Branch on `auth_type`.
2. For PAT: use `token` as Bearer (REST) or `WorkspaceClient(host=..., token=..., product=..., product_version=...)` (SDK). Use connector-level product/product_version for User-Agent, not from user config.
3. For M2M: obtain token (REST: call `/oidc/v1/token`, cache with expiry; SDK: `Config(host=..., client_id=..., client_secret=..., auth_type="oauth-m2m", product=..., product_version=...)`). Use connector-level product/product_version.
4. For U2M/token: use provided `access_token` as Bearer (REST) or `WorkspaceClient(host=..., token=access_token, ...)` (SDK). If the connector runs the browser flow, do that once and then store the access_token (and optionally refresh_token) in config or session. User-Agent from connector level.

**Rule:** Do not mix auth methods in one process. If the app supports multiple connection types, each connection instance should use only one auth type (see auth isolation in the main cursor rule).

---

## 3. Operations through the client

All API or SDK operations go **through** the client returned by `connect()`:

- **REST:** Every request uses the client’s `Authorization: Bearer <token>`, `User-Agent`, and `Content-Type` where needed. For M2M, the client refreshes the token when expired (same pattern as in rest-api-authentication.md).
- **Python SDK:** Use `client.tables.get()`, `client.jobs.list()`, etc. No extra auth or User-Agent wiring per call.
- **Databricks Connect:** Use `DatabricksSession.builder.sdkConfig(config).getOrCreate()` with a single `Config` built from host, auth (token or client_id/client_secret with `auth_type="oauth-m2m"` or U2M token), and compute (`serverless_compute_id="auto"` or `cluster_id`). Set product/product_version on Config. Do not set both serverless and classic in the same Config.

This keeps telemetry and auth in one place and avoids mistakes (e.g. missing User-Agent or wrong token).

---

## 4. Optional: validate_connection()

To verify credentials after connect, run the **two recommended validation tests**:

1. **UC table by full name:** `GET /api/2.1/unity-catalog/tables/<catalog>.<schema>.<table>` (no warehouse). Or with SDK: `client.tables.get(full_name=...)`.
2. **Statement Execution:** `POST /api/2.0/sql/statements` with `warehouse_id` and a simple SQL (e.g. `SELECT 1` or `DESCRIBE EXTENDED catalog.schema.table AS JSON`). Requires `warehouse_id` in config.

If either fails with 401/403, the connection is invalid or missing permissions. Optionally expose this as `client.validate_connection()` or a standalone helper.

---

## 5. Token refresh and secrets

- **M2M:** Cache the access_token and expiry; before expiry (e.g. 5 minutes), call the token endpoint again and update the cache. See REST rule and integration checklist.
- **U2M:** When the access_token expires, use the refresh_token with the OAuth token endpoint (`grant_type=refresh_token`) to get a new access_token. Do not send the refresh_token on every API request.
- **Secrets:** Never log tokens or secrets. Prefer env vars or a secret store for credentials; see integration checklist (“Tokens never logged”).

---

## 6. Summary

| Layer | Responsibility |
|-------|-----------------|
| **User-Agent** | **Required, connector-level.** Set on every request; format `<isv>_<product>/<version>`. Coded in the connector (not a user-configurable option). |
| **Config** | host, auth_type (recommend PAT + M2M + U2M), credentials for that type (see “User input by auth type” above), optional warehouse_id/http_path. No user-supplied User-Agent. |
| **connect(config)** | Single entry point; branch on auth_type; return REST client or WorkspaceClient with correct auth and User-Agent. |
| **Client** | All operations; consistent headers (REST) or SDK client; M2M token refresh inside client. |
| **validate_connection()** | Optional; run UC table get + Statement Execution to verify. |
| **Token refresh** | M2M: refresh before expiry; U2M: refresh_token → new access_token. Never log tokens. |

**Recommendation:** Support all three auth types (PAT, OAuth M2M, OAuth U2M) so users can choose the right one per connection. User-Agent is required but is set at the connector level (in code), not by the end user; the end user provides only host, credentials for the chosen auth type, and optional warehouse_id/http_path. Following this structure keeps auth isolation, telemetry, and validation aligned with the REST and Python auth rules and the integration checklist.

