# Streams in

> Land a database's tables or an app's records in your warehouse, new and changed ones only. Postgres is ready to use; for an app, cherry writes the connector and you approve it.

A stream in lands records in your warehouse from another system: the orders in your shop's database, the sales from your point of sale, the bookings from your booking system, the jobs from your field service app. Each run reads only what is new or changed since the last, so a stream can run every few minutes and stay cheap.

You find streams in under **Data · Streams**, on the **Streams in** tab, once Streams is switched on for your console. Until then the console doesn't list it.

## Two ways in

- **A Postgres database** is read by a *published source*: code sanda wrote and runs as its own, so every workspace has it and there is nothing to approve. You connect the database, choose its tables, and set a schedule. See [stream a Postgres database](https://docs.sanda-os.com.au/streams/postgres).
- **An app's API** is read by a *connector* cherry writes for your workspace from the app's documentation, and an owner or admin approves. The rest of this section is about connectors.

Press **New stream in** for either. When your workspace has approved connectors, **Reads from** chooses between them and Postgres.

## How a stream in is made from an app

You don't write code. You ask cherry, and you approve what it built.

1. **cherry reads the app's API documentation.** You name the app, and the address of its API reference. cherry reads only sites you named in the conversation, and the pages linked from them on the same site.
2. **cherry writes a connector.** A connector says which addresses to ask for, how to read each answer, and how to find the next page. sanda checks it before saving it: it must name the hosts it talks to, and it may only read.
3. **You connect a credential.** cherry opens a form beside the conversation, and you paste the app's API key there. It never goes through the chat.
4. **cherry tries it.** It reads the first pages of each thing the connector reads, with your credential, and lands nothing. You see the records, the columns they would become and every request.
5. **You approve it.** An owner or admin reads what the connector may touch (its hosts, and what it may do there) and approves it. Then cherry asks to start the stream, on the schedule you want.

See [build a stream in with cherry](https://docs.sanda-os.com.au/streams/build-a-stream-in) for the whole conversation, and [approve a connector](https://docs.sanda-os.com.au/streams/approve-a-connector) for what approving means.

## What a connector can and cannot do

A connector is code cherry wrote from documentation, so sanda treats it as code it cannot trust and gives it nothing to misuse:

- **It has no network of its own.** It returns each request as a description, and sanda makes the request.
- **It talks only to the hosts it declared, and that an owner approved.** A request anywhere else is refused before it is sent, and so is a redirect that leads anywhere else. Hosts that accept data from anybody, such as request catchers and tunnels, are refused outright.
- **It only reads.** It may send GET requests, and POST only to the search addresses it declared, for the APIs whose searches are a POST.
- **It never sees your credential.** sanda attaches the credential to each request itself, and only to requests for the hosts it was connected for.
- **It has no database.** It hands records back, and sanda lands them.
- **It is paced.** At most 20 requests a second, and fewer when the connector says the API wants fewer. When the API asks sanda to slow down, it does.

## Where the records land

Each thing a stream reads (a *resource*, such as sales or customers) lands as a table of its own in your warehouse's `raw` schema, named after the stream and the resource: a stream called "pos" reading sales lands in `raw.pos__sales`.

- **One row per record**, found by the record's key, so a record read twice is still one row.
- **A column per field.** Each field of the record becomes a column, typed by what its values are: numbers, dates, times with their zone, true or false, text, or JSON for a nested object or list. A string of digits, such as an account number or a postcode, stays text.
- **New fields become new columns** when they first arrive with a value, up to 200 columns. A column is never retyped or dropped, so a view built on it keeps working.
- **The whole record is kept** in `_record`, as the API sent it. When a value no longer fits its column (text in a number column, a date that doesn't exist), the column is empty for that row and the value is still in `_record`. The run says how many values didn't fit.
- **`_synced_at`** is when the row last changed, and **`_deleted_at`** is set when a full read no longer finds the record. Nothing is ever deleted.

The tables are landing tables like a connection's, so you put them on your [semantic map](https://docs.sanda-os.com.au/semantic-fluid/the-map) the same way, and **Learn from my data** reads them with the rest of your warehouse.

## New and changed records only

A resource is read one of two ways. A connector declares which, and for a Postgres table you choose ([how each table is read](https://docs.sanda-os.com.au/streams/postgres#how-each-table-is-read)):

| Read | What a run reads | When a record disappears |
|---|---|---|
| **New and changed** | Records updated since the last run, by a field such as `updated_at`. Each run starts a little before the newest record the last one saw, so a record updated while the last run was reading is not missed | It keeps its row: the API never says it went |
| **Every record** | The whole resource, every run. For small lists, such as locations or staff | Its row is marked `_deleted_at` |

Either way, a record that has not changed since sanda last landed it is neither rewritten nor counted.

A large first load is several runs. One run lands at most 5,000,000 records, and the next run carries on from the page the last one reached.

## Load everything again

**Load everything again**, on a stream's page, makes the next run read every record from the start, as its first run did. Records already landed are only counted again if they changed. Use it when an app changed records in a way it doesn't report as an update.

## Credentials

A stream signs in with a credential an owner or admin connected for its connector: an API key, a username and password, or a key the app sends in a header. sanda checks it with the app before keeping it, keeps it encrypted, and sends it only to the hosts it was connected for, which are fixed when it is connected. It is never shown again, in the console or to cherry.

The **Streams in** tab lists each connector's credentials under **Connectors**, so you can tell them apart by the last characters of a key, or by the username a password goes with. **Type it again** replaces the value (after you rotate a key at the app), and **Disconnect** forgets it. A stream whose credential the app stops accepting waits, rather than failing again and again, until someone types it again. A Postgres database's username and password are listed the same way under **Published sources**, with **Sign in again** in place of **Type it again**.

If somebody pastes a key into cherry's chat anyway, sanda takes it out of the message before it is kept or answered, and says so. If it was a real key, change it at the app.

:::links
- [Build a stream in with cherry](https://docs.sanda-os.com.au/streams/build-a-stream-in): The conversation, from the app's name to a running stream.
- [Approve a connector](https://docs.sanda-os.com.au/streams/approve-a-connector): What you are agreeing to, and what changes when a connector is fixed.
- [Runs](https://docs.sanda-os.com.au/streams/runs): What each run did, and what to do when one fails.
:::
