Learn from my data
Ask cherry to read your warehouse and propose tables, metrics, dimensions, relationships and names for you to review. It uses AI, which is billed.
Learn from my data gets you a first model without typing one. cherry reads the shape of your warehouse, asks an AI to describe what it means, and files what comes back as suggestions. Nothing goes live until you accept it, so the worst outcome of a bad pass is a queue you reject.
Before you start
- Rows have landed. Sync a connection or upload a CSV first. A connection that has never run has nothing to read.
- cherry has a model to run on. If the button is greyed out and its tooltip says cherry has no model to run on yet, an owner or admin can check Settings · cherry.
- cherry may work with your data. The data switch under Settings · cherry covers running a learn pass.
- You are an owner, admin or member. cherry changes nothing for a viewer.
- Today's AI limit is not used up. A pass is AI usage. See cost and time.
Run a pass
-
Open the Semantic fluid page. In the console, go to Intelligence · Semantic fluid.
-
Press Ask cherry to learn. cherry's pane opens on a new conversation, and cherry starts work at once.
The same request sits behind Learn from my data on a connection's page (while none of its tables are on the map), on the finish screen of a CSV upload, and in cherry's setup checklist. Those send it in the conversation you are already in.
-
Watch it read. cherry shows a step for each connection, such as "reading Billing (1 of 2)". Once every connection is read, it shows "finding the relationships across your map".
-
Read the card. Under cherry's words, a card headed "learned from N tables" reports what was proposed. cherry adds what stands out and what it could not settle.
-
Review the suggestions. The suggestions panel opens on the map. See Review what sanda proposes.
You can also ask in your own words: "learn from my Billing connection" reads one connection alone.
What a pass reads
For each table, sanda reads a description of its shape:
- The table's name, its columns and their types, and an approximate row count.
- Up to six short sample values per column, taken from a small sample of rows. Each value is shortened.
- A check of the table's key against a sample of its rows, and a measurement of likely joins. These are counts only, never values.
A pass considers at most the 60 largest tables in your warehouse, and up to 120 columns of any one table. Within a wide table it keeps key columns first, then columns you have already named, event times, id-like names and numbers. The AI is told how many columns were left out. The loader's own bookkeeping columns are never read.
If your warehouse holds more than 60 tables, the smallest are left out. Add those yourself from new table on the map. A pass that reads one connection applies the limit to that connection's own tables, so a small connection beside a large one is read in full.
Columns whose names suggest a secret, such as a password, token, API key, card number or tax file number, contribute their type and nothing else.
What it proposes
| Proposal | What you get |
|---|---|
| Tables | A table worth querying, with a name, whether it is a fact or a dimension, its grain, its key columns and its event time column |
| Dimensions | A column worth slicing by. Dates get a grain and a roll-up, and category columns get sample values |
| Metrics | A number defined as SQL over one table, such as sum(gross_amount) - sum(discount_amount) |
| Names | Business names for columns, and value labels for stored codes |
| Relationships | Joins between tables, each with which end holds one row |
| Questions | What only you can settle: the currency, the start of your financial year, whether a status code means what its name says |
Every proposal carries a confidence between 0 and 1. Near 0.9 means the data made it plain, around 0.5 means a reasonable reading of a name, and lower means it is worth a glance. Tables are also filed into their connector's dataset and given a place on the map.
Relationships come from two sources. First sanda proposes the joins your schema states: a column named like another table's key, or a foreign key the source declares. Then it asks the AI about the rest, with the measured evidence in front of it. See Relationships.
A pass does not propose named filters, categories or glue. Add those yourself. Table aliases come from suggest aliases on the map, not from a pass.
How sanda checks it before you see it
- Proposals go through the same checks as a form you fill in by hand. A metric must be a single SQL fragment on one table, and a column that already has a dimension cannot get a second one.
- Each metric is test-planned against its real table, which reads no rows. A metric that does not run against its table, for a column that does not exist say, is dropped with the reason. A metric that builds on another metric cannot be checked this way and is counted as unchecked.
- A relationship that names a column sanda did not show the AI is dropped.
- A join whose names match but whose values do not line up is not proposed. Fewer than half of the sampled values found in the other table's key counts as not lining up.
- Accepted definitions are never overwritten, and a suggestion you rejected is not proposed again under the same name.
The card lists what was dropped under N dropped · show why.
Read the result
The card reports, in this order:
- What was proposed: tables, dimensions, metrics and mappings, and from how many tables. It notes how many were grouped into a dataset and given a place on the map.
- Relationships: how many were drawn from matching keys and how many sanda proposed. If there were none, a suggest relationships button offers another try.
- Metrics that could not be checked, and N dropped · show why for proposals that were dropped.
- sanda could not settle these, the open questions.
- sanda's notes, and any connection the pass did not run for.
When something was proposed, review the suggestions opens the queue.
Cost and time
AI is billed. Every model call in a pass is metered as AI usage, priced by the token, and counted against your daily AI limit. The default limit is A$5 a day, and owners and admins can raise it as far as A$50. A trial's free credit pays for a pass first. See usage and budgets and limits.
A pass makes one AI call for each connection it reads, one after another, then one relationship search across the whole map, which asks the AI about the joins your names did not settle. A workspace with four sources pays for one relationship search, not four. cherry's own messages in the conversation are AI usage too.
It can take a few minutes, and longer with many connections. The result stays in cherry's conversation, so you can reopen it later.
If today's limit is used up, the pass stops and says so. The limit resets at midnight UTC, or an owner or admin can raise it.
Run it again
Run a pass again whenever new data lands. It extends what you have and skips anything already on the map.
The map has smaller AI-assisted actions that use the same reading:
- suggest aliases in the aliases panel proposes the words your team would use, and metrics your data makes measurable.
- suggest relationships in the relationships panel searches for joins.
- describe a metric in the metrics panel drafts a metric from a sentence, as a form you edit.
- Map with AI in the glue dialog proposes column pairs between two datasets. Nothing is saved until you press Create glue.
Each is billed the same way.
If nothing was proposed
- Nothing has landed yet. cherry says so. Sync first.
- Everything is already on the map. cherry says nothing new was proposed.
- The daily limit is used up. The card names the limit and when it resets.
- A connection failed. The card names the connection and the reason, and the other connections still count.
See troubleshooting if a pass keeps failing.
Something unclear or out of date? Tell us, and we will fix the page.