Databricks Metric View → ThoughtSpot Model
Converts a Databricks Unity Catalog Metric View into a ThoughtSpot Model. Reads the
Metric View YAML definition via DESCRIBE TABLE EXTENDED, maps dimensions and measures
to ThoughtSpot columns and formulas, translates SQL expressions, and imports the result
via ts tml import.
Two scenarios are supported:
- Scenario A (existing tables): ThoughtSpot Table objects already exist for the Databricks source table(s) the Metric View references. Reuses those existing Table objects.
- Scenario B (new tables): No ThoughtSpot Table objects exist yet for the Databricks source table(s). Creates new Table objects pointing to those objects.
Ask one question at a time for dependent decisions (each answer narrows the next — target database, then schema, then table). Batch independent questions when possible — e.g. connection name + target database + schema can be collected together (BL-074).
References
| File | Purpose |
|---|---|
| ../../shared/mappings/ts-databricks/ts-from-databricks-rules.md | Databricks MV YAML parsing, type mapping, formula translation, column classification |
| ../../shared/mappings/ts-databricks/ts-databricks-formula-translation.md | SQL → ThoughtSpot formula translation rules (bidirectional reference) |
| ../../shared/mappings/ts-databricks/ts-databricks-properties.md | Property coverage — what maps, what doesn't, what is partially migrated |
| ../../shared/schemas/databricks-metric-view.md | Databricks Metric View YAML schema (v0.1 single-source, v1.1 multi-source) |
| ../../shared/schemas/thoughtspot-tml.md | TML export parsing (PyYAML pitfalls, type detection) |
| ../../shared/schemas/thoughtspot-table-tml.md | Table TML structure, connection reference, data types, import patterns, common errors |
| ../../shared/schemas/thoughtspot-model-tml.md | Model TML structure, join scenarios, formula visibility, self-validation checklist |
| ../../shared/schemas/thoughtspot-sql-view-tml.md | SQL View TML structure — sql_view: type for subquery sources |
| ../../shared/schemas/thoughtspot-formula-patterns.md | ThoughtSpot formula syntax, all function categories, LOD/window patterns, YAML encoding rules |
| ../ts-profile-databricks/SKILL.md | Databricks auth methods, profile config, CLI usage |
| ../ts-profile-thoughtspot/SKILL.md | ThoughtSpot auth methods, profile config, CLI usage |
| ../../shared/worked-examples/databricks/ts-to-databricks.md | End-to-end Dunder Mifflin conversion: LOD dimensions, semi-additive, cross-measure ratios, multi-fact split |
| ../../shared/worked-examples/databricks/ts-from-databricks.md | End-to-end MV → Model conversion: direct + computed dimensions, simple + ratio + window + conditional measures |
| ../../shared/worked-examples/databricks/ts-from-databricks-sql-view.md | Subquery source path: SQL View TML + Model on top |
| references/concept-mapping.md | Full Concept Mapping table — Databricks MV YAML construct → ThoughtSpot Model, row by row |
| references/console-templates.md | Console/report templates for Steps 8A (table plan), 10 (review checkpoint), 10-FILE (file-only report), and 12 (summary report + test questions) |
Concept Mapping
Full Databricks Metric View YAML → ThoughtSpot Model mapping table (33 rows,
covering source:, dimensions, measures, windows, joins, and filter:):
references/concept-mapping.md.
Key structural rules:
column_idmust use the column name from the ThoughtSpot Table TML. Export Table TMLs to confirm — do not assume they match the Metric View column names.- Simple measures (
AGG(col)— one column, one aggregate) →MEASUREcolumn. Complex expressions →formulas[]entry. - In Scenario A, the Table TML already exists — reuse its GUID and column names.
- In Scenario B, create a new Table TML from the Databricks schema introspection.
- v1.1 MVs provide
display_name,comment, andsynonyms— mapdisplay_nameto columnname,commenttodescription, andsynonymstoproperties.synonyms(withproperties.synonym_type: USER_DEFINED). Synonyms at the column root level are silently dropped on import. v0.1 MVs only havenameandexpr. - LOD dimensions (with
AGG() OVER (PARTITION BY ...)) map to ThoughtSpotgroup_aggregate()formulas — always 3 arguments:group_aggregate(expr, {query}, query_filters()). - Cross-measure references (
MEASURE(name),ANY_VALUE(dim)) must be inlined as full expressions during TML import —[name]cross-references fail during import (open-items #4). After import, users can simplify formulas in the ThoughtSpot UI. - Duplicate column_id: when the same physical column appears as both an ATTRIBUTE
dimension and a MEASURE (e.g.,
COUNT(col)), convert the measure to a formula to avoid the "unique column_id" import error (open-items #2).
Prerequisites
ThoughtSpot
- ThoughtSpot Cloud instance, REST API v2 enabled
- User account with
DATAMANAGEMENTorDEVELOPERprivilege — only required for import - Authentication configured — run
/ts-profile-thoughtspotif you haven't already - The
tsCLI installed (pip install -e /path/to/tools/ts-cli)
No ThoughtSpot import access? You can still run this skill in file-only mode — it generates the Table and Model TML files for you to import manually. Select FILE at the Step 10 checkpoint or say "file only" at any point before Step 11.
Databricks
- Databricks workspace with Unity Catalog enabled
- SQL warehouse on the Preview channel — Metric Views are a Preview feature;
Current channel warehouses return
PARSE_SYNTAX_ERROR - Databricks CLI installed and profile configured — run
/ts-profile-databricksif you haven't already - A Databricks connection configured in ThoughtSpot (required for Scenario B table creation)
Step 0 — Overview
On skill invocation, display this plan before doing any work:
ts-convert-from-databricks-mv — convert a Databricks Metric View into a ThoughtSpot Model, translating dimensions, measures, and SQL expressions.
Steps:
- Authenticate (ThoughtSpot + Databricks) .............. auto
- List Metric Views in the catalog ..................... auto
- Select a Metric View ................................. you choose
- Fetch the Metric View definition ..................... auto
- Parse the YAML (dimensions, measures, filter) ........ auto
- Map to ThoughtSpot columns + translate expressions .... auto
- Table registration question (reuse or create) ........ you choose
- Discover / create ThoughtSpot Table objects ........... auto (may ask for clarification)
- Build Table + Model TML — ts databricks build-model .... auto 9.5. Confirm Spotter enablement (re-run build-model + flag) . you choose
- Review checkpoint — inspect TML before import ......... you confirm
- Import Model TML — ts databricks build-model --profile . auto
- Verify import and produce summary report .............. auto
File-only mode: at Step 10, choose FILE to write TML files for manual import.
Confirmation required: Steps 3, 7, 9.5, 10 Auto-executed: all others
Ready to start? [Y / N]
Do not begin Step 1 until the user confirms.
Workflow
Step 1: Authenticate
Session continuity: If profiles were already confirmed earlier in this conversation (e.g. for a previous Metric View), skip this step and reuse them.
ThoughtSpot profile:
- Run
ts profiles listto show configured profiles. - If multiple profiles: display a numbered list and ask the user to select one.
- If exactly one profile: display it and confirm before proceeding.
- Verify:
ts auth whoami --profile {name}— print display_name and base URL.
Databricks profile:
Load the profile from ~/.claude/databricks-profiles.json:
import json, os
profiles_path = os.path.expanduser("~/.claude/databricks-profiles.json")
with open(profiles_path) as f:
profiles = json.load(f)
# Display profiles for selection
for i, p in enumerate(profiles, 1):
print(f" {i}. {p['name']} ({p.get('host', '')})")
profile = next(p for p in profiles if p["name"] == "{profile_name}")
dbx_profile = profile["dbx_profile"]
catalog = profile.get("default_catalog", "")
warehouse_path = profile.get("sql_warehouse_http_path", "")
warehouse_id = warehouse_path.rstrip("/").split("/")[-1] if warehouse_path else ""
Verify connectivity:
databricks auth describe --profile {dbx_profile}
Store dbx_profile, catalog, and warehouse_id for use in subsequent steps.
SQL execution pattern
All Databricks SQL in this skill uses the Statement Execution API:
databricks api post /api/2.0/sql/statements \
--profile {dbx_profile} \
--json '{"warehouse_id": "{warehouse_id}", "statement": "{sql}", "wait_timeout": "50s"}'
The response contains:
{
"status": {"state": "SUCCEEDED"},
"manifest": {"schema": {"columns": [...]}},
"result": {"data_array": [[...], ...]}
}
If status.state is PENDING, poll the statement ID:
databricks api get /api/2.0/sql/statements/{statement_id} --profile {dbx_profile}
Step 2: List Metric Views
Query the catalog for available Metric Views:
SELECT table_catalog, table_schema, table_name
FROM system.information_schema.tables
WHERE table_type = 'METRIC_VIEW'
AND table_catalog = '{catalog}'
Execute via the SQL execution pattern above. Display results as a numbered list:
Metric Views in {catalog}:
1. {schema}.{view_name_1}
2. {schema}.{view_name_2}
...
If no Metric Views are found, check:
- The catalog name is correct
- The SQL warehouse is on the Preview channel
- The user's role has
USE CATALOGandUSE SCHEMAgrants
Step 3: Select a Metric View
If the user has already named a Metric View, skip this step.
Otherwise, ask the user to select from the list displayed in Step 2 (enter a number
or type a fully qualified catalog.schema.view_name directly).
Store the selected {catalog}, {schema}, and {view_name} for subsequent steps.
Step 4: Fetch the Metric View definition
Retrieve the YAML definition via DESCRIBE TABLE EXTENDED:
DESCRIBE TABLE EXTENDED {catalog}.{schema}.{view_name}
Execute via the SQL execution pattern. Parse the response result.data_array — each
row is [col_name, data_type, comment, metadata].
Extract the definition:
- Find the row where
col_name == 'View Text'— thedata_typecolumn contains the full YAML string. - Find the row where
col_name == 'Type'— confirmdata_type == 'METRIC_VIEW'. - Store the YAML string for parsing in Step 5.
If the query fails with "table or view not found", verify the fully-qualified name
and confirm the user's role has SELECT on the view.
Step 5: Parse the YAML — ts databricks parse-mv
Parse the YAML string extracted in Step 4 with the deterministic parser (schema: ../../shared/schemas/databricks-metric-view.md):
printf '%s' "$MV_YAML" > mv.yaml
ts databricks parse-mv mv.yaml --output parsed.json
The command handles version routing (0.1 / 1.1), fields:/dimensions: aliasing,
dimension/measure classification (including window: and its 5 range values),
the nested joins: walk (on/using, cardinality/rely precedence), the
materialization: block, and metadata pass-through. It prints WARNING: lines
(e.g. BL-098 density checks) on stderr.
Exit 1 with unsupported[]: the JSON is still written; each entry names the
construct and why. Show the list to the user and ask whether to skip this MV
(log to the Unmapped Report) or handle the construct manually. Never continue
silently past a non-empty unsupported[].
Read parsed.json and branch on source.kind:
sql_query— the source is a SELECT subquery. Present the (D / T / M / S) options below.table_fqnwithneeds_live_check: true— run the MV-on-MV live check below before treating it as a physical table. The same check applies to everyjoins[].source.
Present these options to the user:
| Option | Action |
|---|---|
| (D) Create Databricks VIEW | Execute CREATE VIEW catalog.schema.view_name AS {source} in Databricks, then use the new view as a regular table FQN. Re-enter Step 5 with the new FQN. |
| (T) Create ThoughtSpot SQL View | Build a sql_view: TML that runs the source SQL against the Databricks connection, import it, then build a Model on top. See Step 5T below. |
| (M) Map to existing | User provides an existing ThoughtSpot Table/View name — skip table creation, proceed to Model. |
| (S) Skip | Log the MV to the Unmapped Report and continue to the next MV. |
If the user selects (T), proceed to Step 5T. Otherwise continue with the selected option. For (D), (M), and (S), the rest of Step 5 proceeds as normal (or skips).
Step 5T — Build a ThoughtSpot SQL View from a subquery source:
See ../../shared/schemas/thoughtspot-sql-view-tml.md for the full TML reference.
Read
source.rawandfilterfromparsed.json, then construct the SQL query:sql_query = parsed["source"]["raw"] # the SELECT statement if parsed.get("filter"): sql_query = f"SELECT * FROM ({parsed['source']['raw']}) _mv WHERE {parsed['filter']}"Introspect the query columns. Execute the query with
LIMIT 0via the Statement Execution API to retrieve the column schema without returning data:SELECT * FROM ({sql_query}) _cols LIMIT 0Parse
manifest.schema.columnsfrom the response — each entry hasnameandtype_name.Build the
sql_view:TML:sql_view: name: "{model_name}" description: "SQL View for Databricks MV {mv_fqn}. Source: {source_fqn}" connection: name: "{databricks_connection_name}" sql_query: "{sql_query}" sql_view_columns: - name: "{col_name}" sql_output_column: "{col_name}" properties: column_type: ATTRIBUTE # or MEASURE for numeric columns used in aggregationImport via
ts tml import --profile {profile} --policy PARTIAL --create-new. Record the returned GUID.Continue to Step 6 (mapping) — treat the SQL View as the source table. Column references in the Model TML use
[SQL_VIEW_NAME::column]syntax. The Model'smodel_tablesentry references the SQL View name and GUID.
See ../../shared/worked-examples/databricks/ts-from-databricks-sql-view.md for a complete worked example.
MV-on-MV detection (when source.kind is table_fqn): an FQN-shaped source:
cannot be assumed to be a physical table — it may be another metric view. Query:
SELECT table_type FROM system.information_schema.tables
WHERE table_catalog = '{src_catalog}' AND table_schema = '{src_schema}'
AND table_name = '{src_table}'
If table_type = 'METRIC_VIEW', fail loud — do not build a Table TML against it.
Report to the user:
The Metric View's source ('{source_fqn}') is itself a Metric View, not a physical
table. Chained (MV-on-MV) sources are not supported by this skill yet. Convert
'{source_fqn}' on its own first, or flatten the source chain in Databricks before
retrying.
Log the MV to the Unmapped Report and stop processing it. This same check applies
to joins[].source values.
Version routing and classification are handled by parse-mv — both versions
normalize into one parsed.json shape.
See ../../shared/mappings/ts-databricks/ts-from-databricks-rules.md
(v1.1 Multi-Source Parsing section) for background reading on the join-walk
semantics parse-mv implements.
Step 6: Translate expressions — ts databricks translate-formulas
1. Author tables.json — the alias-path → ThoughtSpot-table-name map:
- Key
"source"(required) plus one key per join alias; nested joins use dot-joined paths from the root ("orders.customers"). - Value: the ThoughtSpot Table object name. Derive the initial name from the
FQN's table part upper-cased (
analytics.ecommerce.transactions→TRANSACTIONS); Step 8 confirms it against the real objects and enriches these values with GUIDs.
{"source": "TRANSACTIONS", "orders": "DM_ORDER", "orders.customers": "DM_CUSTOMER"}
2. Translate:
ts databricks translate-formulas \
--input parsed.json --tables tables.json --output translated.json
The command encodes the full mapping surface — the function map from
../../shared/mappings/ts-databricks/ts-databricks-formula-translation.md,
the window decision tree from
../../shared/mappings/ts-databricks/ts-from-databricks-rules.md,
conditional aggregates, LOD group_aggregate (3-arg with query_filters()),
COUNT(DISTINCT) → unique count, and cross-measure inlining. Every
dimension/measure lands in translated[] or skipped[] with a reason —
review both.
3. Review the output with the user:
skipped[]— each entry needs a decision: accept the omission, or build the formula manually per the reason text (e.g.range: allneeds a partition-dimension judgment call — the reason names the manual recipe).annotations[]— surface everysparse_data_risk(BL-098: Databricks date-interval windows vs ThoughtSpot row-positionalmoving_sum; numbers match only on data dense at the window grain),pending_verification(C8 grain/offset extrapolations),one_row_per_period, andlod_filter_asymmetrynote. These feed the Step 10⚠ WINDOWmarkers.- The
filterentry (when present) becomes the model-level filter in Step 9.
Semantic caveats (live-verified):
Density caveat (E1, live-verified 2026-07-09 on gapped data — see
docs/audit/2026-07-09-dbx-semantic-claim-matrix.md). Databricks'
trailing/leading N day windows are genuine date-interval windows;
ThoughtSpot's moving_sum/moving_average are row-positional — the two are
indistinguishable on dense daily data, but diverge on data with date gaps.
Treat any translated[] entry carrying a sparse_data_risk annotation as an
approximation that matches Databricks only when the order column is dense at
the window's unit grain (one row per unit, no gaps); run a density check on
any source with possible gaps before accepting the translation.
Filter-asymmetry caveat (A1/A2/A3, live-verified 2026-07-09 — same matrix).
The range: all / LOD group_aggregate(..., query_filters()) mapping is
filter-aware for a Databricks MV's own global filter:, but filter-blind to
an ad hoc query-time WHERE on an MV with no global filter — unless the
third argument is {} instead of query_filters(), paired with a
model-level filters: block mirroring the MV's filter: (live-tested
2026-07-09, A3 follow-up): that combination is filter-blind to a
search-level pin and filter-aware of the model filter, reproducing both
Databricks conditions in one formula. Every entry carrying a
lod_filter_asymmetry annotation needs this judgment call made explicitly
with the user.
Deferred grains note (C8). The range: current + offset: -N <unit>
row-relative LAG(N) mapping is live-verified at month grain (N=1) only;
quarter/year grain offsets remain Deferred (C8 — see
docs/audit/2026-07-08-dbx-window-claim-matrix.md). Entries carrying a
pending_verification annotation at those grains need this caveat surfaced
to the user before the translation is treated as final.
Step 7: Table registration question
After mapping, display the source table(s) and ask:
v0.1 (single source):
The Metric View references 1 source table:
{catalog}.{schema}.{table_name}
Is this table already registered in ThoughtSpot?
Y Yes — use existing ThoughtSpot Table object
N No — create a new Table object from scratch
? Not sure — search ThoughtSpot first
Enter Y / N / ?:
v1.1 (multi-source with joins):
List all source tables (fact table from source: + all dimension tables from joins:)
and ask the same question for each.
- Y → skip search, go to Step 8A (column verification only)
- N → skip search, go to Step 8B (create)
- ? → go to Step 8A (search + verify)
Step 8A: Discover and verify existing ThoughtSpot Table objects (Y and ? paths)
Skip this step if the user answered N in Step 7 — go directly to Step 8B.
Choose the search scope first. A whole-instance scan is the slow path — on a
large instance --all pulls every table. Offer the narrower option and search by
table-name pattern (--name), never --all-then-filter:
How should I search for these tables?
C Within a specific connection — fastest; search that one connection's tables
I Entire ThoughtSpot instance — broader, slower
Enter C / I :
Search by name (both scopes start here):
ts metadata search --subtype ONE_TO_ONE_LOGICAL --name "%{table_name}%" --profile {profile}
- C (within a connection) → first identify the connection using the
N (name it) / F (filter by substring) / L (list all) prompt in Step 8B — present that
prompt and let the user choose; do NOT run
ts connections listand dump every connection by default. Then pass--connection "{connection_name}"tots metadata searchrather than hand-filtering: the CLI scopes it infilter_by_connection(commands/metadata.py:35), which casefolds both sides. This step previously said to keep results whosedataSourceNameequals the connection name — an executor following that literally drops rows the CLI keeps (APJ_DBXvsapj_dbx). Corrected 2026-08-26, finding 11.1. Fastest, and unambiguous when the same table name exists on several connections. - I (entire instance) → run the name search above with no connection filter.
Filter the JSON to match the MV source table by table name (metadata_name).
Connection scoping is already done by the --connection flag above — do NOT re-filter on
metadata_header.dataSourceName here: the flag casefolds and a hand comparison does not, so
re-applying it drops every row the flag kept (finding 11.1, 2026-08-26). Use
metadata_header.database_stripes / metadata_header.schema_stripes to disambiguate
same-named tables. Build a map: physical_table_name -> {metadata_id, metadata_name}.
Only fall back to
--all(fetch every table) when no usable name pattern can be formed (e.g. the name is too generic). Tell the user that cost before running it.
Export TMLs for found tables to verify columns:
ts tml export {guid1} {guid2} ... --profile {profile} --parse
--parse returns structured JSON — access columns via item["tml"]["table"]["columns"]
directly. Parse table.columns[].name from each returned item. Build a column map:
table_name -> [col_name, ...]. Compare against the columns referenced in the MV
dimensions and measures to identify any column gaps.
The
column_idin the model TML must use the column names from the ThoughtSpot Table TML — export the TMLs to confirm them.
Confirm the plan before making any changes: console template — references/console-templates.md "Step 8A — Table plan confirmation console template".
Do not proceed until the user confirms. If any table is not found, follow Step 8B for those tables. If any table has missing columns, follow Step 8C before building the model.
Enrich tables.json: upgrade each value from a name string to an object
carrying the ThoughtSpot GUID, e.g.
{"source": {"name": "TRANSACTIONS", "fqn": "<guid>"}}. Step 9's
build-model stamps these GUIDs as model_tables[].fqn. For file-only
runs (Step 10-FILE) where new tables were needed but NOT created, instead set
"create": true with db/schema/db_table and a columns list of
{"name", "dbx_type"} from the DESCRIBE TABLE output — build-model then
writes a {name}.table.tml alongside the model.
Step 8B: Create ThoughtSpot Table objects (Scenario B) — also the connection picker for the Step 8A connection-scoped search
Get all column names and types for the source table:
DESCRIBE TABLE {catalog}.{schema}.{table_name}
Execute via the SQL execution pattern. The response data_array contains rows
[col_name, data_type, comment]. Map Databricks types to ThoughtSpot types using
the data type mapping in
../../shared/mappings/ts-databricks/ts-from-databricks-rules.md.
Choose the ThoughtSpot connection using the N/F/L connection selection flow in
../../shared/references/connection-select.md,
with --type DATABRICKS on the ts connections list call. No E/C prompt — Databricks
connection creation is out of this skill's scope (see note below).
No suitable connection? A ThoughtSpot connection only sees catalogs its Databricks credentials are granted. If no existing connection can see the source catalog, table creation fails with "Database … does not exist in connection". Creating a Databricks connection is out of this skill's scope — Databricks connections authenticate with a Personal Access Token or OAuth (M2M), not key-pair, so there is no
ts connections createpath for them (that isBL-036). Stop and direct the user to create the connection in the ThoughtSpot UI (credentials per/ts-profile-databricks), then resume on this step and select it. Do not trial-and-error existing connections.
Create the ThoughtSpot Table object:
cat tables-spec.json | ts tables create --profile {profile}
Where tables-spec.json is a JSON array built from the column data. See
ts tables create --help for the spec format. This command handles JDBC retry and
GUID resolution automatically, and outputs {name: guid}.
Both
dbandschemamust be non-empty in every table spec entry. ThoughtSpot requires the fully qualified three-part path (catalog.schema.table) even if the connection has a default catalog configured. Passing an emptydborschemacauses "Fully qualified table mapping missing".
Record the created GUID for use in the model TML.
Enrich tables.json: upgrade each value from a name string to an object
carrying the ThoughtSpot GUID, e.g.
{"source": {"name": "TRANSACTIONS", "fqn": "<guid>"}}. Step 9's
build-model stamps these GUIDs as model_tables[].fqn. For file-only
runs (Step 10-FILE) where new tables were needed but NOT created, instead set
"create": true with db/schema/db_table and a columns list of
{"name", "dbx_type"} from the DESCRIBE TABLE output — build-model then
writes a {name}.table.tml alongside the model.
Step 8C: Update existing tables with missing columns
For each table from Step 8A with a column gap, introspect the Databricks schema for the missing columns:
DESCRIBE TABLE {catalog}.{schema}.{table_name}
Filter the result to the missing column names. Map Databricks types to ThoughtSpot types using ../../shared/mappings/ts-databricks/ts-from-databricks-rules.md.
Find the ThoughtSpot connection for the table:
ts connections list --type DATABRICKS --profile {profile}
Add the missing columns to the connection, then re-import the updated Table TML (batch all imports in one call):
ts tml import --policy ALL_OR_NONE --profile {profile}
After import, re-export the updated TMLs to refresh the column map before Step 9.
Step 9: Build the Model TML — ts databricks build-model
Model name: {view_name_title_case} — derived from the Metric View name.
Ask the user if they want a different name. Do not add a TEST_MV_ or other
prefix — see ../../shared/schemas/ts-model-conversion-invariants.md (N1).
ts databricks build-model \
--parsed parsed.json --translated translated.json --tables tables.json \
--connection "{connection_name}" --model-name "{model_name}" \
--mv-fqn "{catalog}.{schema}.{view_name}" --output-dir ./tml_out
The command assembles the Model TML (columns, formulas, joins, the model-level
filters: block when the MV has a filter:, properties:), assembles Table
TML for any create: true entries, validates the hard TML invariants
(db_column_name on every table column, name:-only connection blocks,
guid: at document root, formula/column formula_id pairing, no
aggregation: in formulas[]) and runs the ts tml lint checks (I1/I2/I4/
I5/I8) — exit 1 names the specific field on any finding. It prints a summary
JSON on stdout; keep it for Step 10.
Filter handling is automatic: the MV filter: becomes an MV Filter
ATTRIBUTE formula plus a model-level filters: block (oper: in,
values: ['true']) — the live-verified emulation of the MV's global filter
(A1, docs/audit/2026-07-09-dbx-semantic-claim-matrix.md). If the filter was
skipped at translate time it appears in skipped[] — resolve it with the user
before importing (an unfiltered model silently changes every number).
Live-verified 2026-07-09 (see docs/audit/2026-07-09-dbx-semantic-claim-matrix.md,
A1) — this model-level filters: approach is the CONFIRMED correct emulation of
the MV's own global filter:: on ThoughtSpot, a model-level filters: block makes
LOD/window formulas (e.g. group_aggregate(..., query_filters())) filter-aware
identically to a query-level pin, matching the DBX MV's global-filter: condition
exactly. It does not, and cannot, reproduce a DBX consumer's ad hoc query-time
WHERE on an MV with no global filter — that DBX condition is filter-blind for
LOD/window dimensions, a source-side asymmetry, not a ThoughtSpot gap.
build-model emits these correctly; they matter when you hand-edit TML:
idmust equalnameexactly (same case, same characters). ThoughtSpot resolveswithandonjoin references against the table's actualname— ifiddiffers in case, joins fail with "{table_name} does not exist in schema". Use the exact ThoughtSpot table object name for bothidandname.idvalues must be unique across allmodel_tablesentriesnamevalues must also be unique — ThoughtSpot rejects models where two tables share the samenamevalue ("Multiple tables have same alias")
Step 9.5: Spotter enablement
Ask whether Spotter (AI search) should be enabled. Default is yes.
Enable Spotter (AI search) for this model? [Y / n] (default: Y)
Re-run the Step 9 build-model command appending --spotter-enabled (or
--no-spotter-enabled) — assembly is deterministic and cheap; the flag adds
properties.spotter_config.is_spotter_enabled to the model TML. When
updating an existing model (--existing-guid) and the user does not
explicitly answer: export the existing model first (ts tml export {guid} --profile {profile} --parse), read its current
properties.spotter_config.is_spotter_enabled, and pass the matching flag —
never overwrite a live setting with a default.
Step 10: Review checkpoint
Before importing, show the user a summary built from the Step 9 build-model summary
JSON (and translated.json/parsed.json for the formula log): tables from tables[]
(alias/name/fqn), column lists from columns.attributes/columns.measures,
formula translations from translated.json's translated[]/skipped[] entries,
window markers from window_measures[] (name, ts_expr, each annotation's detail
as the assumption line), and Spotter from spotter_enabled. Console template:
references/console-templates.md
"Step 10 — Review checkpoint console template".
Wait for user confirmation before proceeding.
If the user selects file, skip to Step 10-FILE.
Step 10-FILE: Output TML files (file-only mode)
This path is used when the user selected file at the Step 10 checkpoint, explicitly
said "file only", or has no ThoughtSpot DATAMANAGEMENT access.
1. Files are already on disk — Step 9's build-model --output-dir wrote
{model_name}.model.tml (and {table_name}.table.tml for every
create: true table). Nothing to generate here.
2. Report: console template — references/console-templates.md "Step 10-FILE — File-only mode report template".
3. Proceed to Step 12 (Produce summary report) — include the formula translation log and column summary so the user has the full picture before importing.
Pre-import validation gate
Before any ts tml import, run the mandatory lint gate — see
../../shared/schemas/ts-tml-import-gate.md
for the invariant list, the stdin command, and the
update-vs-create guid and import-policy rules. Do not import until
ts tml lint reports "clean": true.
Step 11: Import the model
IMPORTANT — Updating vs creating: without --existing-guid, ThoughtSpot
always creates a new object, even when a model with the same name exists.
To update in place, pass the existing model's GUID — build-model stamps it
as a top-level guid: alongside model: (a guid: nested inside model: is
silently ignored). On first import omit it; record the returned GUID.
# First import (new model):
ts databricks build-model \
--parsed parsed.json --translated translated.json --tables tables.json \
--connection "{connection_name}" --model-name "{model_name}" \
--mv-fqn "{catalog}.{schema}.{view_name}" --output-dir ./tml_out \
{--spotter-enabled|--no-spotter-enabled} --profile {profile}
# Update an existing model in place:
# ... same command plus: --existing-guid {existing_model_guid}
With --profile, the command imports the model TML via
ts tml import --policy PARTIAL after a clean lint, and reports
import_status and model_guid in the summary JSON. Save the GUID —
required for any future update import. On import_status: "failed" read
import_error and consult the common error table.
Common import errors: see
ts-tml-import-gate.md § 4.
Step 11b: Verify Import
Follow ts-tml-import-gate.md § 5.
Step 12: Produce summary report
After a successful import (or file output), generate the Markdown report (Model Import Complete header, Columns Imported, Formula Translation Log, Not Mapped) and, separately, 3-5 suggested Spotter test questions based on the dimensions and measures present. Both templates: references/console-templates.md "Step 12 — Summary report template" / "Suggested test questions template".
Step 13: Cleanup
Remove any temporary files written during the workflow:
rm -f /tmp/ts_model_build_*.yaml /tmp/ts_model_build_*.json
The ts CLI manages its own token cache — do not remove /tmp/ts_token_*.txt
unless the user explicitly requests a logout.
Multiple Metric View conversion
After completing Step 12 for one view, ask: "Convert another Metric View?" If yes: return to Step 2. Reuse the already-confirmed ThoughtSpot and Databricks profiles. Do not re-authenticate between views.
Changelog
| Version | Date | Summary |
|---|---|---|
| 1.13.0 | 2026-09-02 | BL-232 — column descriptions reached the TML at the wrong nesting level and were silently discarded on import (ts-cli v0.136.0). An MV's comment: was written to columns[].properties.description, but ThoughtSpot expects description as a sibling of name and a Model import silently ignores unknown keys inside properties — so the TML linted clean, imported with status_code OK, and the descriptions were gone. Live-caught 2026-09-02 converting dunder_mifflin_sales_mv: 19 of 19 descriptions sent, 0 stored; relocating them to the column root and re-importing the same GUID restored all 19. Synonyms were unaffected because they were already placed correctly two lines away. This mattered most for Spotter, which reads column descriptions as AI context. The reverse leg (build-mv) had the mirror-image bug — it read properties.description, so a genuine ThoughtSpot Model's descriptions never reached an emitted MV comment:; the two cancelled out in a TS→MV→TS round-trip, which is how both survived. Worked examples corrected. |
| 1.12.2 | 2026-08-26 | Use ts metadata search --connection instead of hand-filtering dataSourceName; the old instruction said equals where the CLI casefolds (finding 11.1). |
| 1.12.1 | 2026-08-26 | Carry BL-074's prompt-batching rule — ask one question at a time for dependent decisions, batch independent ones. The rule reached 13 skills but omitted the four conversion skills, which are the most interactive in the repo by ask-count (finding 14.6). A check_patterns rule now enforces it above a question-count threshold. |
| 1.12.0 | 2026-07-31 | BL-174 — generated Model joins get LEFT OUTER semantics, and a measure's format: currency now reaches the TML (ts-cli v0.127.0). (1) joins[].type was INNER on every join, unconditionally. A Metric View has no join-type field because Databricks fixes it — "In a star schema, the source is the fact table and joins with one or more dimension tables using a LEFT OUTER JOIN" (Joins in metric views, re-confirmed 2026-07-31) — so every fact row with a NULL or unmatched FK was kept by the MV and dropped by the Model, making measures read lower in ThoughtSpot than in Databricks with nothing in the pipeline warning. Now LEFT_OUTER, nested joins included (live-validated on se-thoughtspot, VALIDATE_ONLY). This changes the numbers a converted Model returns — re-convert any Model whose fact FKs are nullable. (2) format: {type: currency, currency_code: USD} on a measure now writes properties.currency_type.iso_code; the data reached translated.json and was discarded at assembly, though the reverse leg already implemented the pair. type: percentage, decimal_places and a format: on a fields:/dimensions: entry stay unmapped — declared as new limitation L12. (3) joins[].cardinality keeps being written: omitting it is refused by the platform ("both type and cardinality should be defined", l |
…(truncated)