ENTERPRISE DATA PLATFORMS

Christian Støa

Enterprise Data and Analytics Leader

I build the governed data platforms, decision intelligence, and AI enablement that let complex organizations run on facts. Fifteen years bridging operations, finance, and technology, currently leading enterprise data science and platform strategy at Velocity Vehicle Group, a $4B multi-entity organization spanning eight global business units.

15+
Years across ops, finance, tech
$4B
Multi-entity org scale
400+
Reports governed
8
Global business units

Selected work

Enterprise data infrastructure built from real operating data. The two diagrams below document a platform running in production today.

Estate Telemetry Pipeline

A self-running pipeline that measures an entire Power BI and Fabric estate, from capture through a Direct Lake semantic model with a deployed AI data agent. Built so that silence is the failure mode and every figure stays traceable by someone who did not build it.

Estate Telemetry Pipeline Measuring the Power BI & Fabric estate · From daily manual capture to a pipeline that runs itself VIEW E · TELEMETRY Audience: Enterprise Data Team As at 1 October 2026 LAYER 1 Telemetry Sources LAYER 2 Capture & Orchestration LAYER 3 plat_landing · Lakehouse LAYER 4 plat_warehouse · Silver + Gold LAYER 5 Model & Decisions GONE IN DAYS · CAPTURE OR LOSE rolling Admin activity log shallow window Documented window loses the newest day. Fabric Capacity Metrics shallow window CU, duration, throttle. Refresh history API shallow, per model A busy model exhausts it in under two days. Usage metrics, legacy longer window Page grain, client telemetry. Usage metrics, new shorter window Event grain, server telemetry. WHY RETENTION IS THE WHOLE POINT A documented window of N days yields N−1 complete days, because a calendar day is asked for at 00:00Z. Proven three times. Nothing rebuilds a day that fell out, at any price, so every run backfills the window. DURABLE · SNAPSHOT OR CORPUS Scanner API weekly Workspaces, items, tables, measures. TMDL definitions corpus append-only Append-only. One folder per workspace. TWO USAGE MODELS, TWO MEASURES Legacy counts every page change. The new model counts one report open, and page views as a separate measure. Measured against each other, they diverge widely, so a shared label would silently double-count. They join on grain, never on the word “views”. DAILY · 08:00 UTC · ONE PIPELINE runMultiple nb_10_activity backfill nb_20_capacity restated nb_40_pageviews legacy nb_41_usage_new events Window from the watermark, never “today minus N”. THEN, SAME PIPELINE · T-SQL Script activities gold.usp_rebuild 18 tables gold.usp_rebuild_usage 7 tables gold.usp_snapshot_counts row census gold.vw_rebuild_health OK or PROBLEMS Script activities, not a notebook. A gold failure fails the run. EVERY SIX HOURS nb_30_refresh no day window A shallow per-model window burns through quickly. WEEKLY · SUNDAY 09:00 UTC nb_50_scan estate shape nb_90_propose_refs → ref_*, a person confirms ONE TIME nb_05_seed_refs the judgement nb_99_migrate one-time history load ON DEMAND · SCOPED BY DOMAIN no schedule nb_61_definitions TMDL → Files nb_62_model_audit three parsers, unchanged nb_63_audit_load findings → bronze Written 1 October, not yet run. The first pull is the proof. STILL ON THE WORKSTATION what is left estate_state.py assembles the published basis enrich_sql_export.py owes sql_sources.csv, so bronze is empty THE ALARM IS SILENCE Two verdicts leave the daily run: the notebooks, then gold. No complete day in 48 hours fails the schedule. The audit is deliberately outside that check, having no window to fall behind. BRONZE · 18 TABLES · AS COLLECTED bronze_ omitted activity activity_event event capacity capacity_item_day item · day capacity_day day refresh refresh_operation operation refresh_attempt attempt refresh_schedule item · capture usage, legacy page_view_day page · day page_view_model coverage usage, new report_view_event open report_page_view_event page usage_report dim usage_model_new coverage estate scan_workspace scan scan_item scan audit_run run audit_finding run × item × check sql_source empty definition_export the TMDL manifest LEGACY · 3 TABLES · FROZEN AT CUTOVER legacy_refresh_daily frozen at cutover legacy_view_daily frozen at cutover legacy_view_user_daily partial REF · 3 TABLES · THE JUDGEMENT, AS DATA ref_workspace domain · environment · production ref_operation what counts as use ref_account person or service OPS · 2 TABLES · DID IT ACTUALLY RUN ops_watermark last complete day per source ops_load_run one row per run, pass or fail HOW EVERY WRITE LANDS A replace over a date partition, or a MERGE on a natural key. nb_90 proposes new workspaces and accounts into ref_* · a person confirms. WHAT HAS LANDED Event-grain usage now lands cleanly: report opens, page views and distinct people, with service accounts separated out. Consumption splits across browser, export, mobile, embedded, PowerPoint and Outlook. None unclassified. Legacy remains the only record before cutover. SILVER · VIEWS ONLY · NO TABLES 32 views USE vw_activity events + judgement vw_view_daily_all Fabric + legacy vw_view_coverage what each day can answer REFRESH PRESSURE vw_refresh_run operation grain vw_refresh_pressure duty on rebuild time vw_refresh_failure failed refreshes COST AGAINST READERSHIP vw_capacity_day CU per item per day vw_readership cost vs readers vw_page_usage with coverage attached SHAPE & HEALTH vw_estate_item what exists, by domain vw_audit_trend findings over time vw_pipeline_health did every source land USAGE · EVENT GRAIN vw_report_view_event one open, classified vw_report_view_daily report · day · class vw_report_page_view_daily page grain USAGE · DISTRIBUTION & EXPORT vw_report_app_usage distribution_method vw_external_tool XMLA, with the pipeline flagged vw_export_profile scheduled or a person GOLD · TABLES, NOT VIEWS · THE MODEL BASIS 25 tables · rebuilt nightly NAME MAPS gold.map_workspace one key per spelling gold.map_item one key per spelling DIMENSIONS · 7 dim_date dim_workspace dim_item dim_operation dim_account dim_consumption dim_dataset FACTS · 16 fact_activity fact_view_daily fact_page_view_daily fact_capacity_daily fact_refresh_operation fact_refresh_attempt fact_refresh_schedule fact_estate_item fact_audit_finding fact_pipeline_run fact_coverage_daily fact_report_view_daily fact_report_page_view_daily fact_report_reader_daily fact_external_tool_daily fact_export_profile WHY TABLES AND NOT VIEWS Direct Lake on OneLake reads the Delta files, so a view is not a source it can use at all. The nightly rebuild is DELETE then INSERT in one transaction, never DROP, so the model's framing survives the night. The create scripts do drop, which is why they are one-time and never part of the schedule. THE JOIN RULE, MADE STRUCTURAL The normalising expression exists twice, both inside the rebuild procedures. Every fact joins a raw name to a map to get its key, so “join on name_key(), never on the raw strings” holds by the shape of the model rather than by everyone remembering to do it. THE CUTOVER IS DERIVED The boundary between legacy and Fabric is read from the data rather than typed in, so it moves by itself after a backfill. fact_coverage_daily carries the same idea forward: readers is NULL where never measured rather than 0, and has_attempt_detail flags the history gap. SEMANTIC LAYER Direct Lake on OneLake Built 1 October · 23 tables · 47 relationships No import. No refresh. No concurrency slot consumed on a capacity where slots are the binding constraint. It reads the Delta files, so the SQL endpoint is not in the path and any restriction is RLS in the model. PUBLISHED · PRODUCTION BASIS Power BI estate: current state Council-facing · every figure traceable Estate Decisions Register What was decided, and what was rejected Estate Capture Runbook 28 failure modes · about half fail silently DECISIONS THIS PIPELINE FEEDS What to retire Cost against readers, with the grain declared Where the capacity ceiling is Refresh slots. Low duty % can still hide a full queue Which models to restage Direct Lake removes the refresh entirely Which reports are extracts Non-browser consumption, by report WHAT THIS STILL CANNOT SAY Isolated days · no activity events exist, at any price Before cutover · report opens do not exist, only legacy page changes, which count something else Before capacity metrics landed · throttle, rejected and failed are unknown rather than zero Page names · SectionId only, until reports are scanned Workspace count · scanned vs published, still open Findings by domain · gold still sums across audit runs, which double counts a model covered twice SQL dependency · no rows, so no page External-only models · no report-to-dataset binding Stated as gaps, because a confident wrong number costs more than an honest one. GOVERNANCE COUNCIL The measurement is the evidence base. Access, retirement and capacity decisions rest on it. A person confirms every ref_* proposal. BUILDING ON IT WITHOUT BREAKING IT One owner of the model. Reports connect live, and Desktop never imports, because an import makes a second estate. The model now holds the measures. Where it and the build sheet disagree the sheet is wrong, in three logged places. IDENTITY, CONTROL AND THE THINGS THAT STOP IT RUNNING Service principal Named identity, survives a role change kv-platform-secrets Tenant, client id and secret. None in source Read-only admin APIs No admin-consent permission, of either kind Workspace access Usage metrics reach only what the identity sees ops_watermark · ops_load_run Replaces somebody remembering it ran Silence is the failure mode Alert on no complete day in 48 hours WHY THIS EXISTS Retention Windows close nightly · a missed day is gone at any price Grain Attempts, not operations · the failure was invisible above it One definition Domain, environment and name spelling live in one place Reproducible Every figure traceable, by someone who did not build it Unattended No daily manual run · no single workstation Honest Gaps declared, never filled with a zero Enterprise Data Platform · Estate Telemetry Pipeline Microsoft Fabric · every object name read from the SQL, every figure from the pipeline that produced it Measure it, or manage it by anecdote. View E · Telemetry

Estate Telemetry Schema

The full schema lineage, Bronze through the semantic model, with the grain of every object declared. Five architectural decisions, each answered by something running in production rather than a preference.

Estate Telemetry Schema Bronze to Silver to Gold to the semantic model · what each layer is for, and the grain of every object VIEW F · SCHEMA Audience: Enterprise Data Team As at 1 October 2026 ACTIVITY & EDITS CAPACITY REFRESH USAGE, LEGACY USAGE, EVENT GRAIN ESTATE & AUDIT BRONZE · plat_landing as collected, never reinterpreted · 18 tables activity_event one event capacity_item_day item × day capacity_day day refresh_operation operation refresh_attempt attempt refresh_schedule item × capture legacy_refresh_daily frozen at cutover page_view_day page × day page_view_model coverage legacy_view_daily frozen at cutover legacy_view_user_daily partial report_view_event one open report_page_view_event one page view usage_report report dim usage_model_new coverage scan_workspace scan × workspace scan_item scan × item definition_export item · the TMDL manifest audit_run run · one scope audit_finding run × item × check sql_source run × item × object REFERENCE & OPERATIONS · the judgement, as data, and the record that it ran ref_workspace domain · environment · is_production ref_operation what counts as use ref_account person or service ops_watermark last complete day per source ops_load_run one row per run, pass or fail SILVER · views only, no tables interpretation · 32 views · one copy of every figure vw_activity event + ref_operation vw_edit_event who changed what vw_capacity_day CU per item per day vw_readership cost against readers vw_refresh_run operation grain vw_refresh_pressure duty on rebuild time vw_refresh_failure failed refreshes vw_refresh_concurrency starts per UTC slot vw_refresh_cutover boundary, derived vw_refresh_daily_all Fabric + legacy vw_view_daily_legacy report × day vw_view_readers_legacy person vw_view_cutover boundary, derived vw_view_daily_all Fabric + legacy vw_view_coverage what a day can answer vw_page_usage with coverage attached vw_workspace_id_name id ↔ name vw_report_view_event one open, classified vw_report_view_daily report × day × class vw_report_page_view_event one page view vw_report_page_view_daily page × day vw_report_app_usage distribution_method vw_external_tool XMLA, pipeline flagged vw_external_tool_daily dataset × day vw_export_profile scheduled or a person vw_usage_coverage_new per workspace vw_estate_item what exists, by domain vw_audit_trend findings over time CONFORMED DIMENSIONS & CROSS-CUTTING FACTS · every fact in the band below joins these NAME MAPS gold.map_workspace one key per spelling gold.map_item one key per spelling DIMENSIONS dim_date dim_workspace dim_item dim_operation dim_account dim_consumption dim_dataset CROSS-CUTTING FACTS fact_pipeline_run run fact_coverage_daily date × source THE RULE Every fact joins a raw name to a map to get its key. The normalising expression lives inside the rebuild procedures and nowhere else, so “join on name_key(), never on the raw strings” holds by the shape of the model rather than by everyone remembering to do it. GOLD · plat_warehouse, physical tables conformed and aggregated · 25 tables · rebuilt nightly fact_activity date × workspace × item × operation × user fact_capacity_daily date × workspace × item fact_refresh_operation request_id · one per operation fact_refresh_attempt attempt · the grain the failure lives at fact_refresh_schedule capture × item · newest capture only fact_view_daily date × workspace × item fact_page_view_daily date × report × page · no item_key fact_report_view_daily date × workspace × item × report × class fact_report_page_view_daily the same, plus page fact_report_reader_daily the same, plus person fact_external_tool_daily date × dataset fact_export_profile as-of × account × report fact_estate_item scan × item fact_audit_finding audit × item × check NOT YET IN GOLD · named rather than drawn, because a planned table on this diagram reads as one that exists fact_sql_source bronze_sql_source holds nothing. enrich_sql_export.py writes an xlsx and owes the CSV, and it is the last thing on the workstation. fact_cross_domain nb_62 writes cross_domain.json and nothing loads it. No landing table, so the cross-domain disagreements are unreportable. findings at the newest run per item Gold sums findings across runs. Once a model is covered by an estate run and a later domain run, it counts twice. report → dataset binding The scan records a report and a model, never which model a report reads. A model read only outside Power BI is the list a retirement decision on view counts gets wrong, and it cannot be produced until the scan carries the binding. SEMANTIC MODEL · Direct Lake on OneLake built 1 October · 23 tables · 47 relationships · 72 measures · 134 columns hidden fact_activity Activity Events · Use Events Model Edits · Unclassified Share the only fact dim_account relates to. Unclassified Share is on the page because RequestCopilot alone dominates the event volume fact_capacity_daily CU Seconds · CU Hours Throttle Minutes · Rejected Operations Harm Detail Available fact_refresh_operation fact_refresh_attempt fact_refresh_schedule Refresh Operations · Attempt Coverage Capacity Rejections · Peak Hourly Starts No fact-to-fact relationship. The attempt carries the operation's refresh_type, status and terminal flag as columns, because it has its own date and its own conformed keys already fact_view_daily fact_page_view_daily Legacy Views · Legacy Page Views fact_page_view_daily has no item_key because the legacy source never named the item fact_report_view_daily fact_report_page_view_daily fact_report_reader_daily fact_external_tool_daily fact_export_profile Report Opens · Page Views Distinct People · App Share Reader Days is not people. Person grain exists because the summed distinct count ran far too high. No edge joins a dataset to the reports that read it, so external-only consumption cannot be listed fact_estate_item fact_audit_finding Workspaces · Production Workspaces Items in Latest Scan · Items Scanned Audit Findings · Measures in Models loaded on demand, per domain, never on the capture schedule THE FIVE OPEN DECISIONS, AS BUILT Each one is answered by something that is running, not by a preference 1 · KEY STRUCTURE Resolve on an id wherever the source carries one. Map through name_key() only where it does not. Never join a raw string. dim_dataset resolves on item_id because joining on the name alone fanned many models out, one of them badly. 2 · WHICH TABLES LAND IN SILVER None. Silver is 32 views and no tables, so there is one copy of every figure and a definition change takes effect on the next read. Materialisation happens once, at Gold, where Direct Lake requires a physical Delta table. 3 · ENRICHMENT IN SILVER Interpretation, and nothing else. Classification, the service's own vocabulary, judgement against ref_*, the derived cutover, coverage flags. No key assignment, no conforming of names, no aggregation. Those need the map tables, which are Gold. 4 · SCHEMA APPROACH A star, in the lake. Direct Lake on OneLake reads the Delta files, so a view is not a source the model can use and every table it reads has to be physical. Single direction, many-to- one, cross-filter Single throughout. dim_dataset carries two keys and relates to neither, and the one fact-to-fact link came out, because both close a path the engine will not resolve. A child fact copies the parent's attributes down. 5 · IS GOLD SUMMARISATION Partly. Gold is the maps, the conformed dimensions, and facts at a declared grain. The grain is chosen by the question, not by a default of summarising: attempt grain exists because the failure was invisible above it, person grain because a summed distinct count ran far above the real people behind it. The audit facts are the open case, below. Enterprise Data Platform · Estate Telemetry Schema Microsoft Fabric · every object name read from the SQL, every grain from the definition that produced it Declare the grain, and the rest follows. View F · Schema

3PL Pricing Model

A cost-to-serve pricing engine for drayage and warehousing. It prices every lane from its real cost drivers, fuel burned at the lane's blended MPG, driver hours, chassis days, and maintenance per mile, then builds to a target margin. The drayage estimator below uses the same method as the full tool.

Lane inputs

Fixed: 6.0 mpg blended, chassis $17/day across 3 days, maintenance $0.085/mile, overhead and SG&A 32.6%, target EBIT 8%.

Estimated rate per move

Fuel$0
Driver$0
Chassis$0
Maintenance$0
Direct cost$0
Overhead and SG&A$0
Target EBIT$0
Customer rate$0
Built to an 8% EBIT margin by construction.

Open the full pricing tool for multi-lane fleets, warehouse activity-based costing, fuel surcharge tiers, and breakeven analysis.

Power BIAzure SQLGovernance

Enterprise BI Platform

Built and govern a Power BI platform of 400+ reports on a unified Azure SQL warehouse integrating 10 source systems. Established reporting standards that eliminated shadow reporting and created a single source of truth for executive and board decisions, including board and bond rating agency packages supporting capital raises.

FP&AOneStreamFinance

FP&A Built From the Ground Up

Stood up the FP&A function supporting $4B in revenue across Parts, Service, and Vehicle Sales. Served as OneStream Global Admin, designing custom drill-throughs that connect executive reporting to transactional detail. Uncovered and resolved overtime pay calculation errors yielding $1.4M in annual savings.

Financial ModelingBusiness CaseLogistics

Capital Business Cases

Delivered a $4M facility redesign business case projecting $7.9M in return over three years, approved and executed by leadership. Translates operational data into capital investment decisions leadership can commit to.

What I build

Not dashboards. Governed data infrastructure, AI-enabled analytics, and decision systems that replace instinct with defensible numbers.

Data platform architecture

Medallion Architecture, Bronze to Silver to Gold, on Microsoft Fabric and Azure SQL. End to end from raw ingestion to certified, query-ready Gold-layer datasets across 10 source systems.

AI and data science enablement

Microsoft Fabric Data Agents deployed against certified Gold-layer data, so natural-language answers rest on trusted numbers rather than whatever a model can scrape.

Governance and master data

Certification gates, a Two-Key approval model, conformed dimensions, and semantic-layer standards that hold a single source of truth across business units.

Decision intelligence

Moving the organization from reporting to governed, actionable insight that leaders can act on directly.

Financial modeling

Driver-based models, cost-to-serve analysis, and pricing intelligence, built on an FP&A function supporting $4B in revenue.

About

Fifteen years building the data platforms, governance frameworks, and decision intelligence that make complex organizations operate more intelligently.

Currently Director of Data Science at Velocity Vehicle Group, a $4B multi-entity organization spanning eight global business units in North America and Australia. Designing and operating Medallion Architecture on Microsoft Fabric, governing the enterprise data layer, and building the AI enablement foundation the analytics run on.

A background across operations, finance, and technology gives a rare ability to connect technical architecture to business outcomes. Earlier roles span Toll Group, Ryder Integrated Logistics, and Dresser-Rand, now Siemens Energy, in Norway.

Microsoft FabricAzure SQLMedallion ArchitecturePower BIData GovernanceAI Data AgentsOneStreamPythonSQLGenAIFP&ADriver-Based Modeling

Director of Data Science

Velocity Vehicle Group · Jul 2026 to Present

Director of Business Intelligence

Velocity Vehicle Group · Nov 2025 to Jun 2026

Sr. Manager, Finance & Analytics

Velocity Vehicle Group · 2023 to 2025

Commercial Finance Manager

Toll Group · 2018 to 2020

Senior Manager

Ryder Integrated Logistics · 2014 to 2018

M.Sc. Business Analytics · MBA

Merrimack College · Cal State San Bernardino

Contact

Open to enterprise data, analytics, and AI leadership roles. The fastest way to reach me is email or LinkedIn.