Skip to content
sandadocs

CSV / files

Read a file from a web address or cloud storage into one table in your warehouse, and read it again on every sync.

This connector reads one file from a location that stays the same, and reads it again on every sync. If you have a file that will not change, or that you would rather drop in by hand, use Load a CSV instead. It puts the file straight into your warehouse with no connector and no schedule.

Before you start

  • A stable location. sanda reads the same address on every sync, so a fixed link to a file that gets replaced beats an email attachment. Each run reads the whole file, so the file should always hold the current data.
  • Access that matches where it lives. A file on the public web needs nothing but its address, and it must open without signing in. A private file needs one of the other Storage Provider options and the credentials that go with it:
Storage Provider What it asks for
HTTPS: Public Web Nothing. The address must open without signing in
GCS: Google Cloud Storage Service Account JSON, the whole key file pasted in
S3: Amazon Access Key ID and Secret Access Key
AzBlob: Azure Blob Storage Storage Account, plus a SAS Token or a Shared Key
SSH: Secure Shell, SCP: Secure copy protocol, SFTP: Secure File Transfer Protocol Host, User, Password and Port (22 unless you changed it)
  • Read-only credentials. For a bucket, create credentials that can read only that bucket. sanda only ever reads.
  • sanda's address allowed, for a file server or a bucket that restricts access by address. See Allowlist sanda's address.

Connect CSV / files

  1. Choose CSV / files. Go to Data · Connections, press New connection and pick CSV / files.
  2. Name the table. Dataset Name becomes the name of the table the file lands in. Use letters, numbers, dashes and underscores only.
  3. Give the address. Put the full address in URL, including its scheme. https://example.com/data.csv, gs://my-bucket/data.csv and s3://my-bucket/data.csv are the shapes the form expects.
  4. Set the format and location if they are not the defaults. Under Advanced settings, File Format opens on csv and Storage Provider opens on HTTPS: Public Web. Change them to match your file, and fill in the credentials the provider asks for.
  5. Continue. Press Continue, name the connection, and press Read schema. sanda opens the file and reads its columns. If it fails, the message says why. See Troubleshoot connections.

What syncs

One file becomes one stream, and lands in your warehouse as one table named after the connection and Dataset Name. File Format accepts csv, json, jsonl, excel, excel_binary, fwf, feather, parquet and yaml. The form warns that some formats may be experimental.

A file carries no record of what changed, so each run reads it in full, and the stream uses a full refresh mode. The default, Full refresh · Overwrite, replaces the table on every run. See sync modes.

Tips

  • Use Reader Options for an unusual layout. Reader Options takes a JSON object of settings for the chosen format. For a semicolon separated file, enter {"sep": ";"}.
  • Header row. By default the first row is read as the column names. If your file has none, set Reader Options to {"header": null}, or the first record is lost.
  • Many files, or a folder? Use a connector that reads by pattern: S3 / object storage, Google Cloud Storage, Azure Blob Storage or SFTP.
  • A spreadsheet rather than a file? Google Sheets reads each tab as a table.

Configuration fields

What the connection form asks for. Required fields are marked; the rest are optional.

Field Type Required Description
Dataset Name string Yes The Name of the final table to replicate this file into (should include letters, numbers dash and underscores only).
File Format string Yes The Format of the file which should be replicated (Warning: some formats may be experimental). One of csv, json, jsonl, excel, excel_binary, fwf, feather, parquet, yaml. Default csv.
Storage Provider one of 8 options Yes The storage Provider or Location of the file(s) which should be replicated.
Reader Options string No This should be a string in JSON format. It depends on the chosen file format to provide additional options and tune its behavior.
URL string Yes The URL path to access the file which should be replicated.

Storage Provider

The storage Provider or Location of the file(s) which should be replicated.

Choose one of the following. Each asks for its own fields.

HTTPS: Public Web

Field Type Required Description
User-Agent boolean No Add User-Agent to request. Default false.

GCS: Google Cloud Storage

Field Type Required Description
Service Account JSON string, secret No In order to access private Buckets stored on Google Cloud, this connector would need a service account json credentials with the proper permissions. Please generate the credentials.json file and copy/paste its content to this field (expecting JSON formats). If accessing publicly available data, this field is not necessary.

S3: Amazon

Field Type Required Description
Access Key ID string No In order to access private Buckets stored on Amazon S3, this connector would need credentials with the proper permissions. If accessing publicly available data, this field is not necessary.
Secret Access Key string, secret No In order to access private Buckets stored on Amazon S3, this connector would need credentials with the proper permissions. If accessing publicly available data, this field is not necessary.

AzBlob: Azure Blob Storage

Field Type Required Description
SAS Token string, secret No To access Azure Blob Storage, this connector would need credentials with the proper permissions. One option is a SAS (Shared Access Signature) token. If accessing publicly available data, this field is not necessary.
Shared Key string, secret No To access Azure Blob Storage, this connector would need credentials with the proper permissions. One option is a storage account shared key (aka account key or access key). If accessing publicly available data, this field is not necessary.
Storage Account string Yes The globally unique name of the storage account that the desired blob sits within.

SSH: Secure Shell

Field Type Required Description
Host string Yes
Password string, secret No
Port string No Default 22.
User string Yes

SCP: Secure copy protocol

Field Type Required Description
Host string Yes
Password string, secret No
Port string No Default 22.
User string Yes

SFTP: Secure File Transfer Protocol

Field Type Required Description
Host string Yes
Password string, secret No
Port string No Default 22.
User string Yes

Local Filesystem (limited)

No further fields.

Something unclear or out of date? Tell us, and we will fix the page.