CSV / files
Read a file from a web address or cloud storage into one table in your warehouse, and read it again on every sync.
This connector reads one file from a location that stays the same, and reads it again on every sync. If you have a file that will not change, or that you would rather drop in by hand, use Load a CSV instead. It puts the file straight into your warehouse with no connector and no schedule.
Before you start
- A stable location. sanda reads the same address on every sync, so a fixed link to a file that gets replaced beats an email attachment. Each run reads the whole file, so the file should always hold the current data.
- Access that matches where it lives. A file on the public web needs nothing but its address, and it must open without signing in. A private file needs one of the other Storage Provider options and the credentials that go with it:
| Storage Provider | What it asks for |
|---|---|
| HTTPS: Public Web | Nothing. The address must open without signing in |
| GCS: Google Cloud Storage | Service Account JSON, the whole key file pasted in |
| S3: Amazon | Access Key ID and Secret Access Key |
| AzBlob: Azure Blob Storage | Storage Account, plus a SAS Token or a Shared Key |
| SSH: Secure Shell, SCP: Secure copy protocol, SFTP: Secure File Transfer Protocol | Host, User, Password and Port (22 unless you changed it) |
- Read-only credentials. For a bucket, create credentials that can read only that bucket. sanda only ever reads.
- sanda's address allowed, for a file server or a bucket that restricts access by address. See Allowlist sanda's address.
Connect CSV / files
- Choose CSV / files. Go to Data · Connections, press New connection and pick CSV / files.
- Name the table. Dataset Name becomes the name of the table the file lands in. Use letters, numbers, dashes and underscores only.
- Give the address. Put the full address in URL, including its scheme.
https://example.com/data.csv,gs://my-bucket/data.csvands3://my-bucket/data.csvare the shapes the form expects. - Set the format and location if they are not the defaults. Under Advanced settings, File Format opens on
csvand Storage Provider opens on HTTPS: Public Web. Change them to match your file, and fill in the credentials the provider asks for. - Continue. Press Continue, name the connection, and press Read schema. sanda opens the file and reads its columns. If it fails, the message says why. See Troubleshoot connections.
What syncs
One file becomes one stream, and lands in your warehouse as one table named after the connection and Dataset Name. File Format accepts csv, json, jsonl, excel, excel_binary, fwf, feather, parquet and yaml. The form warns that some formats may be experimental.
A file carries no record of what changed, so each run reads it in full, and the stream uses a full refresh mode. The default, Full refresh · Overwrite, replaces the table on every run. See sync modes.
Tips
- Use Reader Options for an unusual layout. Reader Options takes a JSON object of settings for the chosen format. For a semicolon separated file, enter
{"sep": ";"}. - Header row. By default the first row is read as the column names. If your file has none, set Reader Options to
{"header": null}, or the first record is lost. - Many files, or a folder? Use a connector that reads by pattern: S3 / object storage, Google Cloud Storage, Azure Blob Storage or SFTP.
- A spreadsheet rather than a file? Google Sheets reads each tab as a table.
Configuration fields
What the connection form asks for. Required fields are marked; the rest are optional.
| Field | Type | Required | Description |
|---|---|---|---|
| Dataset Name | string | Yes | The Name of the final table to replicate this file into (should include letters, numbers dash and underscores only). |
| File Format | string | Yes | The Format of the file which should be replicated (Warning: some formats may be experimental). One of csv, json, jsonl, excel, excel_binary, fwf, feather, parquet, yaml. Default csv. |
| Storage Provider | one of 8 options | Yes | The storage Provider or Location of the file(s) which should be replicated. |
| Reader Options | string | No | This should be a string in JSON format. It depends on the chosen file format to provide additional options and tune its behavior. |
| URL | string | Yes | The URL path to access the file which should be replicated. |
Storage Provider
The storage Provider or Location of the file(s) which should be replicated.
Choose one of the following. Each asks for its own fields.
HTTPS: Public Web
| Field | Type | Required | Description |
|---|---|---|---|
| User-Agent | boolean | No | Add User-Agent to request. Default false. |
GCS: Google Cloud Storage
| Field | Type | Required | Description |
|---|---|---|---|
| Service Account JSON | string, secret | No | In order to access private Buckets stored on Google Cloud, this connector would need a service account json credentials with the proper permissions. Please generate the credentials.json file and copy/paste its content to this field (expecting JSON formats). If accessing publicly available data, this field is not necessary. |
S3: Amazon
| Field | Type | Required | Description |
|---|---|---|---|
| Access Key ID | string | No | In order to access private Buckets stored on Amazon S3, this connector would need credentials with the proper permissions. If accessing publicly available data, this field is not necessary. |
| Secret Access Key | string, secret | No | In order to access private Buckets stored on Amazon S3, this connector would need credentials with the proper permissions. If accessing publicly available data, this field is not necessary. |
AzBlob: Azure Blob Storage
| Field | Type | Required | Description |
|---|---|---|---|
| SAS Token | string, secret | No | To access Azure Blob Storage, this connector would need credentials with the proper permissions. One option is a SAS (Shared Access Signature) token. If accessing publicly available data, this field is not necessary. |
| Shared Key | string, secret | No | To access Azure Blob Storage, this connector would need credentials with the proper permissions. One option is a storage account shared key (aka account key or access key). If accessing publicly available data, this field is not necessary. |
| Storage Account | string | Yes | The globally unique name of the storage account that the desired blob sits within. |
SSH: Secure Shell
| Field | Type | Required | Description |
|---|---|---|---|
| Host | string | Yes | |
| Password | string, secret | No | |
| Port | string | No | Default 22. |
| User | string | Yes |
SCP: Secure copy protocol
| Field | Type | Required | Description |
|---|---|---|---|
| Host | string | Yes | |
| Password | string, secret | No | |
| Port | string | No | Default 22. |
| User | string | Yes |
SFTP: Secure File Transfer Protocol
| Field | Type | Required | Description |
|---|---|---|---|
| Host | string | Yes | |
| Password | string, secret | No | |
| Port | string | No | Default 22. |
| User | string | Yes |
Local Filesystem (limited)
No further fields.
Related
Something unclear or out of date? Tell us, and we will fix the page.