canonic.yaml is validated against CanonicConfig at load time. Unknown top-level fields are ignored rather than rejected. Only version and project are required: everything else defaults.
Top-level fields
version and project are the only fields canonic actually needs to boot. Everything else configures a specific subsystem and can be left out entirely for a local, single-user setup: connections is empty until you add a data source, llm is null until you configure a provider, and the rest fall back to safe local defaults. embeddings and telemetry rarely need touching, they exist mostly for opt-in tuning and privacy control rather than day-to-day project setup.
Environment variables in values
Any value incanonic.yaml can be written as env:VAR, for example host: env:DB_HOST or port: env:DB_PORT. canonic replaces the whole value with the variable’s content when it loads the file, so one committed config can serve dev, staging and production. Only whole values are substituted, there is no ${VAR} syntax and no default value. An unset or empty variable stops the load with an error that names the exact location, such as connections[0].params.host.
Keys ending in _ref (credentials_ref, api_key_ref, token_ref and so on) are not substituted at load time. They stay secret references and are resolved only when the secret is actually needed. canonic setup keeps the env:VAR spelling when it edits an existing file, so resolved values are never written back. The diagnostic bundle from canonic audit --bundle shows the resolved values, so a bug report says which environment it came from.
connections[]
A connection is how canonic reaches an actual data source.
id is the name every other file uses to point back at it, type picks which connector implementation handles it, and params carries whatever that connector needs to connect, host and port for a warehouse, a manifest path for dbt. credentials_ref keeps the secret itself out of canonic.yaml entirely, instead of a password you write a pointer to where the password actually lives.
credentials_refmust be a reference, one ofenv:,keyring:,file:, orprovider:, never a literal secret. Config validation rejects a literal outright.llm.api_key_refand the MCP auth refs accept the same schemes minusprovider:, see Dynamic credentials below.- Postgres/Redshift connections additionally recognize
schema/schemas(search path) andtables(glob patterns narrowing introspection), plusfetch_column_stats: trueto merge zero-scan cardinality/null-ratio stats into the returned schema (a no-op on SQLite/DuckDB). mysqlconnections requirehostanduser. Optional:port(default 3306),database(dbnameis an alias, and without it every database the user can see is introspected),schemas/tables(narrowing introspection, as above),ssl(trueverifies the server certificate against the system trust store),row_limitandstatement_timeout_ms.credentials_refresolves to the password. A MySQL schema is a database, so relations are nameddatabase.table. The driver is not installed with canonic, install it withpip install 'canonic[mysql]'.clickhouseconnections requirehost. Optional:port(default 8123, or 8443 whensecureis set),user(defaultdefault),database(dbnameis an alias),secure(trueconnects over HTTPS and verifies the server certificate),schemas/tables(narrowing introspection, as above),row_limitandstatement_timeout_ms.credentials_refresolves to the password. A ClickHouse schema is a database, so relations are nameddatabase.table. canonic uses the HTTP(S) interface, not the native port. The driver is not installed with canonic, install it withpip install 'canonic[clickhouse]'.snowflakeconnections requireaccountanduser, anddatabasefor introspection. Optional:warehouse,schema,role(overridden byread_only_role),schemas/tables(narrowing introspection, as above),row_limitandstatement_timeout_ms.credentials_refresolves to the password. For key-pair auth setprivate_key_pathto a PEM file, andcredentials_refthen only holds the optional key passphrase. The driver is not installed with canonic, install it withpip install 'canonic[snowflake]'.databricksconnections requireserver_hostnameandhttp_path, andcatalogfor introspection. Optional:schema,schemas/tables(narrowing introspection, as above),row_limitandstatement_timeout_ms.credentials_refresolves to the access token.read_only_roleis not used, because Databricks has no role to switch to. The driver is not installed with canonic, install it withpip install 'canonic[databricks]'.dbtconnections recognizemanifest_path(path to the compiledmanifest.json, defaultmanifest.json) andtarget_connection(the id of the physical, queryable connection whose tables this manifest describes, see Connectors).target_connectionmust name a connection actually declared inconnections[]. An unknown id fails config validation at load time. Omitting it falls back to the dbt connection’s own id.
ossieconnections requirepaths(a list of Apache Ossie model files or glob patterns, relative to the project root, a single string is accepted too) andtarget_connection. Unlike dbt,target_connectionhas no fallback. It must name a declared connection whose type is queryable (postgres,sqlite,duckdband so on), otherwise config validation fails at load time. Supported spec versions are0.1.x(rootsemantic_modelarray) and0.2.0.dev0(one flat model per file). A file with any other version failscanonic connection test, andcanonic ingestingests nothing from it. URL sources are not supported yet.
Dynamic credentials
env:, keyring: and file: all resolve once, to a fixed string that stays valid until someone rotates it by hand. That does not fit a credential the cloud provider issues with a built-in expiry. A Redshift IAM credential lives roughly 15 to 60 minutes, so a connection string built when the daemon starts is dead well before the daemon is.
provider:<name> covers that case. The name selects a registered credential provider, which fetches a fresh credential from the issuer and reports when it expires. Canonic caches it and refetches shortly before it runs out, so a long-lived daemon keeps working without an operator touching anything.
canonic.yaml. cluster_id, db_user and region are not secrets, they tell the provider how to fetch a credential, the same way a file: path is not a secret.
Things worth knowing before you reach for it:
- Redshift and Snowflake only, today.
provider:refs are resolved on every connect, and Redshift and Snowflake are the only connectors that do that so far. The only provider that ships,aws-iam-redshift, issues Redshift credentials, so a Snowflake connection has no provider to point at yet. Pointing any other connector at aprovider:ref fails with a clear error rather than silently freezing the first credential it ever saw. - Not for inbound auth.
llm.api_key_ref,mcp.auth.tokens[].token_refandmcp.auth.oauth.client_secret_refare resolved once at startup and held. A provider credential there would go stale rather than refresh, so config validation rejects it. aws-iam-redshiftneeds boto3, which canonic does not install.pip install boto3. Credentials for the AWS call itself come from the standard boto3 chain (instance role,AWS_PROFILE, environment).- Rejected under
air_gapped: true. Fetching from AWS is network egress, which is what that mode exists to prevent. - Nothing is written to disk. A fetched credential lives in memory for the life of the process that fetched it. A CLI invocation and a running daemon each fetch their own.
llm
Each provider has a fixed credential requirement, enforced at load.
openai_compatible requires base_url, openai and anthropic require api_key_ref, and github_copilot forbids one entirely since it authenticates through a device-code flow that never touches canonic.yaml. Getting this wrong fails config validation immediately rather than surfacing as a runtime error the first time canonic tries to call the model.
reconcile
reconcile only governs how eagerly canonic ingest is allowed to write changes back to semantics/. It has no effect on what evidence gets collected, only on whether a resulting proposal can be applied without a human looking at it first. See Ingestion & reconciliation for how proposals are generated in the first place.
feedback
See Feedback loop (E11) for how these thresholds are used.
telemetry
canonic audit --telemetry-send only sends when all of enabled, endpoint, and transport_acknowledged are set, and runtime.air_gapped is false. Missing any one of them fails closed with a structured error (telemetry_not_configured). Nothing is ever sent implicitly. See Instrumentation & evaluation for the full payload shape and content-safety guarantees.
runtime
When
air_gapped: true, load-time validation additionally rejects a public llm.base_url, telemetry.enabled: true, and any credentials_ref/api_key_ref using a non-local scheme. The check runs once, at config load, so there’s no window where a project starts up in a state that could leak context off the machine. Telemetry has no allowlist path the way llm.base_url does via allow_cidrs. Under air_gapped, it is hard off regardless of telemetry.endpoint/transport_acknowledged.
logging
The daemon writes
.canonic/mcp.log through an inherited file descriptor, so it can only rotate on daemon start, not while the daemon runs. Set logging.file if you need rotation during a long-running daemon.
instrumentation
Rotation moves
served_answer events into a segment. Funnel, outcome, and reconcile events stay in the active file, so limits never affect onboarding state or feedback history. Retention only deletes whole segments and is evaluated when a rotation happens. See The event log.
This applies equally under runtime.air_gapped: true. The log is always local, so nothing else bounds its size.
mcp
This block only matters for
canonic mcp start --transport http. The default stdio transport is a local subprocess and needs no auth, process-level trust is enough. http transport is network-reachable, though, so it refuses to start unless at least one mechanism resolves: a token here (or via the --token-ref CLI override) and/or auth.oauth. When both are configured, a request is checked against the static token map first (no network call), falling through to OAuth verification if no static token matches: the two stay independently revocable, a token entry by editing canonic.yaml, an OAuth-issued token at the IdP. canonic mcp status reports which mechanism(s) are active on a running daemon. See Connecting your agent for the full remote-deployment walkthrough, and Marketplace with Keycloak for a working proxy-mode example against a real IdP.
auth.oauth.verify_id_token matters more than its default suggests. In proxy mode, OIDCProxy verifies the upstream IdP’s access token by default, but many IdPs (Google, GitHub, some Okta setups) issue that as an opaque, non-JWT string, which fails verification outright rather than just producing a poor client identity. Setting verify_id_token: true verifies the id_token instead, which the OIDC spec guarantees is always a standard JWT carrying sub/email claims. This also controls what ends up as client_id in .canonic/events.jsonl: with it off, client_id comes from the access token’s client_id/azp/sub claim (often an opaque subject id). With it on, from the id_token’s, which is far more likely to be a meaningful identity like an email address.Reserved directories
Independent ofcanonic.yaml fields, every project scaffolds four committed context directories plus a git-ignored .canonic/ local-state directory (created with 0o700 permissions):
contracts/policies/ (tenancy.yaml / roles.yaml) is optional. Its presence is the feature switch for tenancy & access control, and a project without it behaves exactly as it did before that layer existed.
raw-sources/<connection-id>/evidence.jsonl
Not a hand-authored config file: it’s a deterministically-sorted, one-JSON-object-per-line snapshot of the evidence ingestion collected for a connection, written by DiskSnapshotStore so re-running ingestion against the same source yields a byte-identical file. Each line validates against EvidenceItem:
This is what
canonic ingest reads to draft proposals against semantics/. You’d only ever read it, not write it.