Running and building demos on your laptop¶
This guide is for sales engineers. It shows you how to start a Provisa demo on your own machine, build a new named demo for a customer scenario, and present it. You need a terminal and Docker. You do not need to write application code.
Everything here runs locally from a checkout of the repository.
Before you start¶
You need:
- Docker Desktop, running. Named demos and
--sourcedemos start containers. - Python 3.12, to create the repository's virtual environment.
- Node.js 20 or later, for the UI. The launcher loads
nvmif you have it and switches to the version in.nvmrc. - Git, to clone the repository.
Set up once:
setup.sh creates .venv and installs the dependencies. The launcher calls .venv/bin/python and .venv/bin/uvicorn directly, so a missing .venv stops it. Run ./venv.sh to rebuild the environment after a pull.
The launcher reads .env in the repository root if one exists, and .env is not committed. Put anything secret there (see Demo hygiene).
Run a demo with one command¶
start-ui-install.sh takes --demo in two forms:
| Command | What starts |
|---|---|
./start-ui-install.sh --demo |
The standard demo: a pet store with four sources |
./start-ui-install.sh --demo <name> |
A named demo from demo/named/<name>/ |
The word after --demo is a name only if it does not start with --. So --demo --source=redis runs the standard demo, and --demo perf --source=redis runs the perf demo.
--demo always runs natively. The backend is a Python process, the federation engine runs inside it, and the control plane is an embedded PostgreSQL database. The standard demo needs no containers.
The standard demo¶
The launcher first resets the demo's control-plane database, so every start is pristine. It then starts a demo petstore API, a demo GraphQL server and the backend, and then the UI. When it finishes it prints:
Provisa running (demo mode):
Backend: http://localhost:8001 (logs: tail -f .../.logs/backend.log)
UI: http://localhost:3000
pgwire: postgresql://admin:ignored@localhost:5439/provisa (username = role)
It also lists the demo sources:
| Source | Type |
|---|---|
pet-store-pg |
PostgreSQL, pet_store schema |
petstore-api |
OpenAPI, http://localhost:18080/api/v3 |
inquiries-sqlite |
SQLite, demo/files/inquiries.sqlite |
graphql-demo |
GraphQL remote, http://localhost:4000/graphql |
Open http://localhost:3000. There is no login: the demo's auth provider is none, and every request acts as the org_admin role. To connect a SQL client over pgwire, use the user name that matches the role you want, for example admin; the password is ignored.
Settings you change in the admin do not survive the next start, except organization settings such as the metadata export target and AI keys. The launcher saves those before the reset and restores them after.
A named demo¶
With a name, the launcher does four more things before it starts the backend:
- Unless you pass
--keep-data, removes the demo's containers and volumes (docker compose ... down -v) and deletesdemo/named/<name>/data, so the data is seeded from empty. Every named-demo start is pristine by default. - Runs
docker compose -f demo/named/<name>/docker-compose.yml up -d --build. - If the compose file defines a service called
seeder, runsdocker compose ... up seederand waits for it. - Uses the demo's own
config.yaml, a whole config, in place of the standard demo config, and printsConfig for named demo '<name>': <path>. (A demo may instead ship afragment.yamlthat the launcher overlays on the standard config; see the pro tip.) A--sourcetoy source is overlaid on either.
With --keep-data, step 1 is skipped. The seeder finds its marker file, prints that the demo is already seeded, and does nothing. Use it for a demo whose data takes too long to regenerate on every start.
perf is a benchmark demo, sized for a server, not a laptop. Its compose file gives PostgreSQL 12 GB of shared buffers and Neo4j an 8 GB heap plus a 4 GB page cache, and it loads 20 million orders, which takes a long time to generate. To run it on a laptop you would need to change two things:
- The row counts. Set them in the environment before you start:
PROVISA_BENCH_ORDERS=200000 PROVISA_BENCH_NEO4J_ORDERS=20000 \
./start-ui-install.sh --demo perf --keep-data
- The memory settings in
demo/named/perf/docker-compose.yml:shared_buffersandeffective_cache_sizefor PostgreSQL, andNEO4J_server_memory_heap_max__sizeandNEO4J_server_memory_pagecache_sizefor Neo4j. Docker Desktop must also be given more memory than the total.
perf takes --keep-data because 20 million rows should not regenerate on every start. The first start also builds a custom PostgreSQL image, which compiles several extensions. For a customer demo, build a demo of your own, as below, and treat perf as a worked example of the layout.
Add a toy source on top¶
See Add one toy source.
Stop and reset¶
- Stop Provisa: press
Ctrl+Cin the terminal that runs the launcher. It stops the backend, the UI and the demo servers. The embedded PostgreSQL stays up for the next start. - Other keys:
Ctrl+Rrestarts the backend,Ctrl+Uclears the UI cache and restarts the UI, andCtrl+Estops the servers and leaves Docker running. - Reset the Provisa side: just start again. Every
--demostart drops and rebuilds the demo control plane. - Reset a named demo's data: just start again. A named demo reseeds from empty by default. Pass
--keep-datato reuse the data from the last seeding instead. - Stop a named demo's data stack.
Ctrl+Cdoes not stop it. Run:
- Force a reseed while using
--keep-data. The seeder skips work whendemo/named/<name>/data/.seededexists. Remove the stack's volumes and the data directory, then start:
docker compose -f demo/named/perf/docker-compose.yml down -v
rm -rf demo/named/perf/data
./start-ui-install.sh --demo perf --keep-data
Deleting data/ alone is not enough for perf: MongoDB keeps its data in a named Docker volume, and down -v removes it. A start without --keep-data does both.
Troubleshooting¶
These are the messages the scripts print, and what to do.
| You see | Cause and fix |
|---|---|
Unknown option: ... followed by a usage block |
A flag the launcher does not know. The valid ones are --demo [name], --keep-data, --source=<name>, --native, --keep-docker, --fast and --idp=basic\|firebase. |
Unknown demo name: X. No demo/named/X/ directory. |
The name does not match a directory. Run ls demo/named. |
--demo X has no config.yaml (and no fragment.yaml) in demo/named/X/ to register it with. |
The demo directory has neither file. Export the model into config.yaml (see Create config.yaml). |
--demo X has both config.yaml and fragment.yaml in demo/named/X/. Keep one: ... |
A demo takes a whole config or a fragment, not both. Delete one. |
Embedded control plane requires pgserver (Python <=3.12). Aborting. |
The virtual environment uses a newer Python. Recreate it with PYTHON=python3.12 ./venv.sh, and read .logs/control-plane-pg.log. |
Backend crashed. Last logs: or Backend did not become healthy. Last logs: |
Read the 20 lines printed, then .logs/backend.log. The launcher waits up to about three minutes for /health. |
UI dev server crashed — see .../.logs/ui.log |
The UI failed to build. Read .logs/ui.log, and check node --version against .nvmrc. |
UI still building after 60s — continuing |
Not an error. Wait, then reload the page. |
Warning: no GitHub token ... GovData subscriptions unavailable. |
Harmless. Only the optional government-data subscriptions need that token. |
Stopping previous start-ui-install.sh instance |
The launcher runs one copy at a time and replaced an older one. |
Docker errors during docker compose up |
The launcher stops at the first failing command. Start Docker Desktop and run the command it printed by hand to see the full error. |
perf demo already seeded (...); skipping |
Normal with --keep-data after the first start. |
rm reports a permission error while removing data/ |
Files written by a container on Linux can belong to another user. Delete demo/named/<name>/data with sudo, then start again. |
Build a new named demo¶
A named demo is a directory, demo/named/<name>/. Create the directory and ./start-ui-install.sh --demo <name> finds it. You do not edit the launcher, and no name is registered anywhere. The launcher refuses a missing directory, and a demo with neither config.yaml nor fragment.yaml, or with both.
The worked example is retail: a PostgreSQL database with customers and orders, plus a stores list loaded from a CSV file.
The files¶
demo/named/retail/
docker-compose.yml the one start point: the database and the seeder
Dockerfile.seeder the seeder's image
seed.py the seeder's entry point
generate_postgres.py creates and fills the tables
stores.csv a small list of stores
config.yaml the demo's whole config: sources, domains, tables, relationships
.gitignore one line: data/
perf has the same shape with four databases instead of one. A demo with no data stores of its own needs only config.yaml: no compose file, no seeder.
1. Create the directory¶
2. Write docker-compose.yml¶
x-demo-labels: &demo-labels
com.provisa.demo: retail
services:
postgresql:
image: postgres:16
restart: unless-stopped
labels:
<<: *demo-labels
com.provisa.demo.role: rdb
environment:
POSTGRES_USER: provisa
POSTGRES_PASSWORD: provisa
POSTGRES_DB: retail
ports:
- "${PROVISA_RETAIL_POSTGRESQL_PORT:-25800}:5432"
volumes:
- ./data/postgresql:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U provisa -d retail"]
interval: 5s
timeout: 5s
retries: 30
start_period: 30s
seeder:
build:
context: .
dockerfile: Dockerfile.seeder
restart: "no"
labels:
<<: *demo-labels
com.provisa.demo.role: seeder
depends_on:
postgresql:
condition: service_healthy
environment:
PROVISA_RETAIL_POSTGRESQL_HOST: postgresql
PROVISA_RETAIL_POSTGRESQL_PORT: "5432"
PROVISA_RETAIL_ORDERS: "${PROVISA_RETAIL_ORDERS:-5000}"
volumes:
- ./data:/data-marker
Four things matter here:
- The service must be named
seeder. The launcher checks for that exact name, and runsdocker compose up seederto wait for it. - Every database service needs a healthcheck. The seeder waits for
service_healthy, so a database that is up but not ready never gets loaded halfway. - Use bind mounts under
./data/for database files. Docker Desktop keeps named volumes inside its own disk image, which fills the system disk. MongoDB is the exception: inperfit uses a named volume, because its storage engine needs file locking that bind mounts on Docker Desktop for Mac do not provide. - Labels.
com.provisa.demogroups every container of the demo, sodocker ps --filter label=com.provisa.demo=retaillists exactly them.
Pick host ports no other demo uses. perf uses 25632, 27317, 28623, 27974, 27987 and 26379. Most toy sources use ports from 21000 to 39999, and Splunk uses 8088 and 8089.
3. Write the seeder¶
Dockerfile.seeder:
FROM python:3.12-slim
RUN pip install --no-cache-dir psycopg2-binary
WORKDIR /app
COPY generate_postgres.py seed.py stores.csv ./
ENTRYPOINT ["python", "seed.py"]
seed.py runs each generator once and records that it did. The marker lives on the ./data bind mount, so it survives container restarts and disappears when you delete data/.
import os
import subprocess
import sys
import time
from pathlib import Path
HERE = Path(__file__).parent
MARKER = Path("/data-marker/.seeded")
ORDERS = os.environ.get("PROVISA_RETAIL_ORDERS", "5000")
def main() -> int:
if MARKER.exists():
print(f"retail demo already seeded ({MARKER.read_text().strip()}); skipping")
return 0
t0 = time.monotonic()
subprocess.run(
[sys.executable, str(HERE / "generate_postgres.py"), "--orders", ORDERS], check=True
)
MARKER.parent.mkdir(parents=True, exist_ok=True)
MARKER.write_text(f"seeded in {time.monotonic() - t0:.0f}s, orders={ORDERS}\n")
return 0
if __name__ == "__main__":
raise SystemExit(main())
generate_postgres.py builds the data from nothing. It reads its host and port from the same environment variables the compose file sets, with localhost and the published port as defaults, so you can also run it from your laptop while debugging.
import argparse
import csv
import os
import random
import time
from datetime import date, timedelta
from pathlib import Path
import psycopg2
from psycopg2.extras import execute_values
HOST = os.environ.get("PROVISA_RETAIL_POSTGRESQL_HOST", "localhost")
PORT = int(os.environ.get("PROVISA_RETAIL_POSTGRESQL_PORT", "25800"))
FIRST = ["Ana", "Ben", "Chi", "Dev", "Eli", "Fay", "Gus", "Hana"]
LAST = ["Ames", "Bose", "Cruz", "Diaz", "Egan", "Frey", "Gray", "Holt"]
STATUSES = ["open", "shipped", "delivered", "returned"]
def connect():
deadline = time.monotonic() + 60
while True:
try:
return psycopg2.connect(
host=HOST, port=PORT, dbname="retail", user="provisa", password="provisa"
)
except psycopg2.OperationalError:
if time.monotonic() > deadline:
raise
time.sleep(2)
def main() -> None:
parser = argparse.ArgumentParser()
parser.add_argument("--orders", type=int, default=5000)
orders = parser.parse_args().orders
rng = random.Random(7) # a fixed seed: every seeding gives the same data
customers = max(orders // 10, 10)
with connect() as conn, conn.cursor() as cur:
cur.execute("DROP TABLE IF EXISTS orders, customers, stores")
cur.execute("CREATE TABLE stores (store_id integer PRIMARY KEY, city text, region text)")
cur.execute(
"CREATE TABLE customers (customer_id integer PRIMARY KEY, name text, email text)"
)
cur.execute(
"CREATE TABLE orders (order_id integer PRIMARY KEY, customer_id integer, "
"store_id integer, order_date date, amount numeric(10,2), status text)"
)
with open(Path(__file__).parent / "stores.csv", newline="") as f:
stores = [(int(r["store_id"]), r["city"], r["region"]) for r in csv.DictReader(f)]
execute_values(cur, "INSERT INTO stores VALUES %s", stores)
execute_values(
cur,
"INSERT INTO customers VALUES %s",
[
(
i,
f"{rng.choice(FIRST)} {rng.choice(LAST)}",
f"customer{i}@example.com",
)
for i in range(1, customers + 1)
],
)
start = date(2025, 1, 1)
execute_values(
cur,
"INSERT INTO orders VALUES %s",
[
(
i,
rng.randint(1, customers),
rng.choice(stores)[0],
start + timedelta(days=rng.randint(0, 364)),
round(rng.lognormvariate(3.7, 0.8), 2),
rng.choice(STATUSES),
)
for i in range(1, orders + 1)
],
)
if __name__ == "__main__":
main()
stores.csv:
The names and addresses come from short fixed lists and @example.com, so no real person appears anywhere.
4. Create config.yaml¶
config.yaml is the demo's whole config: every source, domain, table, relationship and role it shows, plus the settings the standard demo carries. The launcher uses it in place of the standard demo config. Every start rebuilds the control plane from it, so the demo is exactly what the file lists.
You do not write it by hand. Build the demo in the running UI and export it.
- Bring up the data stack yourself, if the demo has one, and start the standard demo (no name):
- In the UI, add the database as a source (PostgreSQL, host
localhost, port25800, databaseretail, userprovisa), register its tables, and add the relationships between them. Add roles or row rules if the story needs them. - Export the model. Open Admin, then Maintenance, then Configuration File, and choose View / Diff. The right pane, Current, is the whole live config. Copy all of it. The same text comes from
curl -s http://localhost:8001/admin/config/livewhile the demo runs. In the demo this export is always on. - Save it as
demo/named/retail/config.yaml. - Add what the export leaves out. It carries your tables, relationships, roles and domains, but not the sources you added in the UI: sources come from the startup file, not from live state. Add each one to the
sources:list by hand, with its connection values as${env:...}references:
sources:
- id: retail-postgresql
type: postgresql
host: ${env:PROVISA_RETAIL_POSTGRESQL_HOST:-localhost}
port: ${env:PROVISA_RETAIL_POSTGRESQL_PORT:-25800}
database: retail
username: provisa
password: ${env:PROVISA_RETAIL_POSTGRESQL_PASSWORD:-provisa}
description: "Retail demo database (PostgreSQL)"
- Decide what else stays. The export starts from the standard demo, so the pet store's sources, tables and relationships are in it. Keep them to show retail beside the pet store, or delete them for a retail-only demo.
Rules that bite:
- Declare
tables:for every source. A baresources:entry never produces tables in the SQL catalog. The export already names each table, its schema, and each column with its type and who may see it. visible_tonames roles. The standard config'sanalystrole may only open the pet-store domains. A new domain is visible toorg_admin, which the demo's requests act as. Add another role underroles:with its owndomain_access, asperfdoes fororg_admin_unguarded.- Connection values are environment variables with defaults:
${env:NAME:-default}. The compose file and the generator read the same names. - Everything is synthetic. Describe it that way in the descriptions.
Pro tip: a fragment instead of a whole config¶
For a small addition to the pet store, you can skip the export and ship demo/named/<name>/fragment.yaml instead of config.yaml. The launcher overlays it on the standard demo config. It holds only list sections, and each is appended to the standard demo's: domains, sources, tables, relationships and roles. A top-level setting the standard config already sets, with a different value, fails the load with config include ...: key '...' conflicts with the including file; only list sections merge.
Give a demo one of the two files. With neither, or with both, the launcher refuses it by name.
The retail demo as a fragment, written by hand:
domains:
- id: retail
description: "Retail demo: orders, customers and stores"
sources:
- id: retail-postgresql
type: postgresql
host: ${env:PROVISA_RETAIL_POSTGRESQL_HOST:-localhost}
port: ${env:PROVISA_RETAIL_POSTGRESQL_PORT:-25800}
database: retail
username: provisa
password: ${env:PROVISA_RETAIL_POSTGRESQL_PASSWORD:-provisa}
description: "Retail demo database (PostgreSQL)"
tables:
- source_id: retail-postgresql
domain_id: retail
schema: public
table: customers
description: Customers (synthetic)
columns:
- {name: customer_id, data_type: integer, is_primary_key: true, visible_to: [org_admin]}
- {name: name, data_type: varchar, visible_to: [org_admin]}
- {name: email, data_type: varchar, visible_to: [org_admin]}
- source_id: retail-postgresql
domain_id: retail
schema: public
table: stores
description: Stores, loaded from stores.csv
columns:
- {name: store_id, data_type: integer, is_primary_key: true, visible_to: [org_admin]}
- {name: city, data_type: varchar, visible_to: [org_admin]}
- {name: region, data_type: varchar, visible_to: [org_admin]}
- source_id: retail-postgresql
domain_id: retail
schema: public
table: orders
description: Orders (synthetic)
enable_aggregates: true
enable_group_by: true
columns:
- {name: order_id, data_type: integer, is_primary_key: true, visible_to: [org_admin]}
- {name: customer_id, data_type: integer, visible_to: [org_admin]}
- {name: store_id, data_type: integer, visible_to: [org_admin]}
- {name: order_date, data_type: date, visible_to: [org_admin]}
- {name: amount, data_type: numeric, visible_to: [org_admin]}
- {name: status, data_type: varchar, visible_to: [org_admin]}
relationships:
- id: retail-orders-customer
source_table_id: orders
source_column: customer_id
target_table_id: customers
target_column: customer_id
cardinality: many-to-one
- id: retail-orders-store
source_table_id: orders
source_column: store_id
target_table_id: stores
target_column: store_id
cardinality: many-to-one
5. Test it¶
Start the data stack by itself first, from the repository root, to check the seeder:
docker compose -f demo/named/retail/docker-compose.yml up -d --build
docker compose -f demo/named/retail/docker-compose.yml logs -f seeder
The seeder exits when it finishes. Confirm the marker and the rows. (The launcher removes this stack on a default start and seeds it again, so this check is only for debugging.)
ls demo/named/retail/data/.seeded
docker compose -f demo/named/retail/docker-compose.yml exec postgresql \
psql -U provisa -d retail -c "select count(*) from orders"
Then run the whole demo:
The launcher tears the data stack down and seeds it again, then uses config.yaml.
In the UI, open Tables and look for the retail domain. Open SQL and run:
SELECT s.region, COUNT(*) AS orders, ROUND(SUM(o.amount), 2) AS revenue
FROM retail.orders o JOIN retail.stores s ON s.store_id = o.store_id
GROUP BY s.region
ORDER BY revenue DESC
If the table names differ in your install, copy the form shown in the Tables list. Then start the demo a second time: without --keep-data it reseeds from empty, and with it the seeder prints retail demo already seeded and skips.
Add one toy source with --source¶
--source=<name> is a shortcut. It starts a source locally in Docker, fills it with sample data, and registers it as a data source, all in one step. Use it when you want to show a specific connector, say Redis or Elasticsearch, without setting up that system yourself.
You could do the same by hand: stand up the system, load some data, and add it as a source in the UI. --source only saves you from standing it up. Each directory under demo/sources/ holds a compose.yml, usually a prime.py that loads a few rows, and a fragment.yaml that registers the source. The launcher starts the containers under the project name provisa-demo-<name>, primes them, and includes the fragment.
./start-ui-install.sh --demo --source=redis
./start-ui-install.sh --demo --source=redis --source=mongodb
List what exists:
Only the sources that also have a fragment.yaml work with --source: cassandra, chinook, elasticsearch, mongodb, neo4j, prometheus, redis, sparql and splunk. For any other directory the containers start and then the launcher stops with --source=<name> has no .../fragment.yaml to register it with.
A --source fragment registers the source only. Add its tables in the UI with Register Table.
An unknown name prints unknown source 'x'; available: ....
To remove a toy source's containers and data:
Which one to use¶
| You want | Use |
|---|---|
| Show one connector, such as Redis or Elasticsearch, next to the pet store | --source |
| A scenario for a customer, with your own tables, relationships and data | A named demo |
| Several databases of different kinds joined together, tables pre-registered | A named demo |
| A few rows and no seeding code | --source |
| Data that must survive restarts, or large data | A named demo, started with --keep-data |
The two combine: ./start-ui-install.sh --demo retail --source=redis.
Demo hygiene¶
- Use synthetic data only. Never copy a customer's data into a demo, not even a sample. Generate rows from fixed lists and a seeded random generator, as
generate_postgres.pydoes. Use@example.comaddresses. - To make a realistic demo from a customer's shape, use profiles and fakes. Profile a table, declare fakes on its identifying columns, and generate a synthetic dataset at the volume you need. Building test environments walks through it, and Fake methods lists every fake. That route copies the distributions and relationships without the people.
- Reset between customer sessions. Start the launcher again (every
--demostart rebuilds the control plane), and a named demo reseeds from empty unless you pass--keep-data. Do not pass--keep-databetween customers if a session changed the data. - Keep secrets out of files. Write connection values as
${env:NAME}inconfig.yamland put the real values in the repository's.env, which git ignores. The launcher loads.envbefore it reads the config. Never commit a password, key or token. Theprovisa/provisapassword in the examples protects a local container that exists only for the demo. - Do not reuse a named demo's ports for anything else running on your laptop.
Presenting it¶
- Open on the Sources page. Show the registered sources and their types. Then show Tables: every source becomes the same kind of table.
- Take the product tour. The first time the UI opens, it offers a guided tour. The compass icon in the toolbar starts it again at any time. It registers a source, exposes tables and then queries them in several languages, in about five minutes. Press
Escor click outside to leave. - Run one query that crosses tables. Use the SQL page. The retail query above joins orders to stores and totals revenue by region. To show a join across sources, join a retail table to a pet-store table, or add a toy source with
--source. - Show the same query in GraphQL. One model, many query languages.
- Show governance. Open Security and show roles and row rules. Connect a SQL client to pgwire (
postgresql://admin:ignored@localhost:5439/provisa) and show the same data there. - Ask Polly. Polly is the data assistant, in a panel at the side of the UI. It needs an LLM key first. Put
ANTHROPIC_API_KEY=...(orOPENAI_API_KEY=...) in your own.envat the repository root before the session. The launcher loads it. Without a key, Polly shows "Polly isn't set up yet" and a button to the AI Models page (/admin/ai-models). For a vendor other than Anthropic, set a model there too. The natural-language page in the tour uses seeded answers and works without one. - End on your customer's question. Name the question first, then show it answered over their sources, in their domain's words.
Before a session, run the launcher once, click through the tour, and ask Polly one question, so nothing surprises you in front of the customer.