diff --git a/_data/navigation.yml b/_data/navigation.yml index e9ebe340b..c06b76681 100644 --- a/_data/navigation.yml +++ b/_data/navigation.yml @@ -659,6 +659,9 @@ items: - url: /storage/tables/ title: Tables & Aliases items: + - url: /storage/tables/incremental-loading/ + title: Primary Keys & Incremental Loading + - url: /storage/tables/data-types/ title: Data Types diff --git a/src/content/docs/catalog/multi-project/index.md b/src/content/docs/catalog/multi-project/index.md index 09a160789..54541ac46 100644 --- a/src/content/docs/catalog/multi-project/index.md +++ b/src/content/docs/catalog/multi-project/index.md @@ -33,7 +33,7 @@ This page explains the concept with a worked example. For a guided walkthrough o Let's say that you have an existing Keboola project that contains: - Oracle database data source connector with the following configurations: - - `ora-history` -- The source database is over 2TB; the largest table is `op_history`, which can be easily loaded [incrementally](/storage/tables/#incremental-loading) as records are only added; it is updated every 10 minutes. + - `ora-history` -- The source database is over 2TB; the largest table is `op_history`, which can be easily loaded [incrementally](/storage/tables/incremental-loading/) as records are only added; it is updated every 10 minutes. - `ora-crm` -- Tables from a CRM system. Major updates are being made to the CRM, so the structure of the tables changes often. Every two weeks, a column is renamed or a table is split into two. - `ora-common` -- Auxiliary tables with addresses and product names that are updated 4 times a year with data from the parent company. - `ora-is` -- Extracts `ORA_IS_XXX` bunch of tables which represent a port of a legacy information system, where column names had to fit into a 6 character limit. diff --git a/src/content/docs/components/extractors/communication/slack/index.md b/src/content/docs/components/extractors/communication/slack/index.md index 35d5dbf72..2058d8d69 100644 --- a/src/content/docs/components/extractors/communication/slack/index.md +++ b/src/content/docs/components/extractors/communication/slack/index.md @@ -16,7 +16,7 @@ Then click **Authorize Account** to [authorize the configuration](/components/#a Select the template you wish to use: -- Smart Mode -- using this mode you always get just missing data (recommended), loads data [incrementally](/storage/tables/#incremental-loading). +- Smart Mode -- using this mode you always get just missing data (recommended), loads data [incrementally](/storage/tables/incremental-loading/). - Full Mode -- using this mode you always get everything. You can download: diff --git a/src/content/docs/components/extractors/database/azure-storage-table/index.md b/src/content/docs/components/extractors/database/azure-storage-table/index.md index 9b08e275c..dcf6b2a6d 100644 --- a/src/content/docs/components/extractors/database/azure-storage-table/index.md +++ b/src/content/docs/components/extractors/database/azure-storage-table/index.md @@ -38,7 +38,7 @@ the [**Configuration Parameters**](#configuration-parameters). Then click **Save - **`table`**: string (required); the name of the input table in the Table storage - **`output`**: string (required); the name of the output CSV file - **`maxTries`**: integer (optional); the max number of retries if an error occurs; the default is `5` -- **`incremental`**: boolean (optional); enables [Incremental Loading](https://help.keboola.com/storage/tables/#incremental-loading); the default is `false` +- **`incremental`**: boolean (optional); enables [Incremental Loading](https://help.keboola.com/storage/tables/incremental-loading/); the default is `false` - **`incrementalFetchingKey`**: string (optional); the name of the key for [incremental fetching](https://help.keboola.com/components/extractors/database/#incremental-fetching) - **`mode`**: enum (optional) - `mapping` (default) diff --git a/src/content/docs/components/extractors/database/cosmosdb/index.md b/src/content/docs/components/extractors/database/cosmosdb/index.md index 833e3d50b..99e4d6b7c 100644 --- a/src/content/docs/components/extractors/database/cosmosdb/index.md +++ b/src/content/docs/components/extractors/database/cosmosdb/index.md @@ -33,7 +33,7 @@ the [**Configuration Parameters**](#configuration-parameters). Then click **Save - **`containerId`**: string (required); the ID of the Cosmos DB container - **`output`**: string (required); the name of the output table in your bucket -- **`incremental`**: boolean (optional); enables [incremental loading](/storage/tables/#incremental-loading); the default is `false` +- **`incremental`**: boolean (optional); enables [incremental loading](/storage/tables/incremental-loading/); the default is `false` - **`incrementalFetchingKey`**: string (optional); the name of the key for [incremental fetching](/components/extractors/database/#incremental-fetching), e.g., `c.id` - **`mode`**: enum (optional) - `mapping` (default) -- items are exported using specified `mapping` diff --git a/src/content/docs/components/extractors/database/index.md b/src/content/docs/components/extractors/database/index.md index 06af4ab09..001e1d7c6 100644 --- a/src/content/docs/components/extractors/database/index.md +++ b/src/content/docs/components/extractors/database/index.md @@ -122,11 +122,11 @@ The rows are fetched from the source table, including the last fetched value. Th ideal to set the ordering column as a primary key so you don't receive duplicated rows in the Storage table. You can clear the stored value if you need to fetch the entire table. -This incremental fetching feature is related to [**incremental loading**](/storage/tables/#incremental-loading). +This incremental fetching feature is related to [**incremental loading**](/storage/tables/incremental-loading/). While not required, it is recommended to turn on incremental loading when fetching data incrementally; otherwise, the table in Storage will contain only newly added rows. This may sound like a good idea when you want to process only the newly added rows. In that case, however, you should do so using -[**incremental processing**](/storage/tables/#incremental-processing). The advantage of using incremental processing +[**incremental processing**](/storage/tables/incremental-loading/#incremental-processing). The advantage of using incremental processing over having only newly added rows in a Storage table is that the table contains all loaded data, and it is not necessary to synchronize extraction and processing. diff --git a/src/content/docs/components/extractors/database/sqldb/index.md b/src/content/docs/components/extractors/database/sqldb/index.md index af04c1ea4..1a492cb13 100644 --- a/src/content/docs/components/extractors/database/sqldb/index.md +++ b/src/content/docs/components/extractors/database/sqldb/index.md @@ -50,10 +50,10 @@ If you want to modify the table extraction setup, click on the corresponding row ![Screenshot - Table Detail](/components/extractors/database/sqldb/sqldb-5.png) You can modify the source table, limit the extraction to specific columns, or change the destination table name in -[Storage](/storage/). The table detail also allows you to define a [**primary key**](/storage/tables/#primary-keys) -and [**incremental loading**](/storage/tables/#incremental-loading). -We highly recommend you define a **primary key** where possible. [Primary keys](/storage/tables/#primary-keys) substantially -speed up the data loads and further table processing. Also, [incremental loading](/storage/tables/#incremental-loading) should be used when possible, which considerably speeds up the data loads. +[Storage](/storage/). The table detail also allows you to define a [**primary key**](/storage/tables/incremental-loading/#primary-keys) +and [**incremental loading**](/storage/tables/incremental-loading/). +We highly recommend you define a **primary key** where possible. [Primary keys](/storage/tables/incremental-loading/#primary-keys) substantially +speed up the data loads and further table processing. Also, [incremental loading](/storage/tables/incremental-loading/) should be used when possible, which considerably speeds up the data loads. Both options require knowledge of the source table, so don't turn them on blindly. ### Advanced Mode diff --git a/src/content/docs/components/extractors/generic-extractor/incremental/index.md b/src/content/docs/components/extractors/generic-extractor/incremental/index.md index dd391bafc..f09317438 100644 --- a/src/content/docs/components/extractors/generic-extractor/incremental/index.md +++ b/src/content/docs/components/extractors/generic-extractor/incremental/index.md @@ -12,7 +12,7 @@ Extracting data incrementally is universally beneficial — it **speeds up the e ## Options After you have incrementally extracted data from an API, the data must be -[incrementally loaded](/storage/tables/#incremental-loading) +[incrementally loaded](/storage/tables/incremental-loading/) into Storage. To do that, simply set `"incrementalOutput": true` in the `config` section. There are, however, a number of implications in the incremental loads. It essentially boils downs to the following use cases, diff --git a/src/content/docs/components/extractors/marketing-sales/bigcommerce/index.md b/src/content/docs/components/extractors/marketing-sales/bigcommerce/index.md index b7853d5ae..d33e1fbdd 100644 --- a/src/content/docs/components/extractors/marketing-sales/bigcommerce/index.md +++ b/src/content/docs/components/extractors/marketing-sales/bigcommerce/index.md @@ -44,7 +44,7 @@ If you are extracting *time-bound* data (i.e., anything **except** `Brands`), yo - **Date From**: only data modified after this date are downloaded. Use the `YYYY-MM-DD` format or a human readable description, e.g., `5 days ago`, `1 month ago`, `yesterday`, etc. You can also set this as `last run`, which will fetch data from the last run of the component; if no previous successful run exists, all data up to specified **Date To** will be downloaded. - **Date To**: only data modified before this date are downloaded. Use the `YYYY-MM-DD` format or a a human readable description, e.g., `5 days ago`, `1 week ago`, `today`, etc. -Finally, in the **Destination** part of the row configuration, you must choose the **Load Type**; i. e., whether you want to use [incremental loading](/storage/tables/#incremental-loading) (by selecting `Incremental Load`) or full loading (by selecting `Full Load`). +Finally, in the **Destination** part of the row configuration, you must choose the **Load Type**; i. e., whether you want to use [incremental loading](/storage/tables/incremental-loading/) (by selecting `Incremental Load`) or full loading (by selecting `Full Load`). ![Row configuration entry](/components/extractors/marketing-sales/bigcommerce/row_config.png) diff --git a/src/content/docs/components/extractors/other/aws-cur-reports/index.md b/src/content/docs/components/extractors/other/aws-cur-reports/index.md index e4da434d8..3375970b6 100644 --- a/src/content/docs/components/extractors/other/aws-cur-reports/index.md +++ b/src/content/docs/components/extractors/other/aws-cur-reports/index.md @@ -55,7 +55,7 @@ and copy the path of the report folder. ![AWS configuration](/components/extractors/other/aws-cur-reports/report_config.png) -If the [**Incremental Load**](/storage/tables/#incremental-loading) is set to true, the new data will be appended to the old ones. +If the [**Incremental Load**](/storage/tables/incremental-loading/) is set to true, the new data will be appended to the old ones. This way you can import new data, e.g., from today, without deleting the data imported before. ## Output Table diff --git a/src/content/docs/components/extractors/other/azure-cost/index.md b/src/content/docs/components/extractors/other/azure-cost/index.md index da15e16db..b45012a7c 100644 --- a/src/content/docs/components/extractors/other/azure-cost/index.md +++ b/src/content/docs/components/extractors/other/azure-cost/index.md @@ -30,7 +30,7 @@ In the [Configuration Row](/components/#configuration-rows) fill in ![Screenshot - Configuration Row](/components/extractors/other/azure-cost/row.png) -If the [**Incremental Load**](/storage/tables/#incremental-loading) is set to true, the new data will be appended to the old ones. +If the [**Incremental Load**](/storage/tables/incremental-loading/) is set to true, the new data will be appended to the old ones. This way you can import new data, e.g., from today, without deleting the data imported before. ## Output Table diff --git a/src/content/docs/components/extractors/other/github/index.md b/src/content/docs/components/extractors/other/github/index.md index 94732775e..835bf6afa 100644 --- a/src/content/docs/components/extractors/other/github/index.md +++ b/src/content/docs/components/extractors/other/github/index.md @@ -15,7 +15,7 @@ to Keboola. Then click **Authorize Account** to [authorize the configuration](/components/#authorization), and select the template you wish to use. There are two configuration templates available: -- `Smart Mode` -- always gets missing data only, loads data [incrementally](/storage/tables/#incremental-loading). +- `Smart Mode` -- always gets missing data only, loads data [incrementally](/storage/tables/incremental-loading/). - `Full Mode` -- always gets everything. ![Screenshot - GitHub configuration](/components/extractors/other/github/github-1.png) diff --git a/src/content/docs/components/extractors/storage/aws-s3/index.md b/src/content/docs/components/extractors/storage/aws-s3/index.md index b9cf5be1c..4289e9e64 100644 --- a/src/content/docs/components/extractors/storage/aws-s3/index.md +++ b/src/content/docs/components/extractors/storage/aws-s3/index.md @@ -143,7 +143,7 @@ The **additional source settings** section allows you to set up the following: - The initial value in **Storage Table Name** is derived from the configuration table name. You can change it at any time; however, the [Storage bucket](/storage/buckets/) where the table will be saved cannot be changed. -- **Incremental Load** will turn on [incremental loading to Storage](/storage/tables/#incremental-loading). The result of the +- **Incremental Load** will turn on [incremental loading to Storage](/storage/tables/incremental-loading/). The result of the incremental load depends on other settings (mainly **Primary Key**). - **Primary Key** can be used to specify the primary key in Storage; it can be used with **Incremental Load** and **New Files Only** to create a configuration that incrementally loads all new files into a table in Storage. diff --git a/src/content/docs/components/extractors/storage/azure-datalake-gen2/index.md b/src/content/docs/components/extractors/storage/azure-datalake-gen2/index.md index b31322eb9..5b8af7cb1 100644 --- a/src/content/docs/components/extractors/storage/azure-datalake-gen2/index.md +++ b/src/content/docs/components/extractors/storage/azure-datalake-gen2/index.md @@ -70,7 +70,7 @@ The **additional source settings** section allows you to set up the following: - The initial value in **Storage Table Name** is derived from the configuration table name. You can change it at any time; however, the [Storage bucket](/storage/buckets/) where the table will be saved cannot be changed. -- **Incremental Load** will turn on [incremental loading to Storage](/storage/tables/#incremental-loading). The result of the +- **Incremental Load** will turn on [incremental loading to Storage](/storage/tables/incremental-loading/). The result of the incremental load depends on other settings (mainly **Primary Key**). - **Primary Key** can be used to specify the primary key in Storage; it can be used with **Incremental Load** and **New Files Only** to create a configuration that incrementally loads all new files into a table in Storage. diff --git a/src/content/docs/components/extractors/storage/ftp/index.md b/src/content/docs/components/extractors/storage/ftp/index.md index d341fe344..81e0f0d4a 100644 --- a/src/content/docs/components/extractors/storage/ftp/index.md +++ b/src/content/docs/components/extractors/storage/ftp/index.md @@ -49,7 +49,7 @@ Now determine how to save the data in Storage. - The initial value in **Table Name** is derived from the configuration table name. You can change it at any time; however, the [Storage bucket](/storage/buckets/) where the table will be saved cannot be changed. -- **Incremental Load** will turn on [incremental loading to Storage](/storage/tables/#incremental-loading). The result of the +- **Incremental Load** will turn on [incremental loading to Storage](/storage/tables/incremental-loading/). The result of the incremental load depends on other settings (mainly **Primary Key**). - **Delimiter** and **Enclosure** specify the CSV settings. diff --git a/src/content/docs/components/extractors/storage/http/index.md b/src/content/docs/components/extractors/storage/http/index.md index 690dd5636..f4155e0bb 100644 --- a/src/content/docs/components/extractors/storage/http/index.md +++ b/src/content/docs/components/extractors/storage/http/index.md @@ -45,7 +45,7 @@ want to load one of our tutorial tables, enter its path, e.g., `/tutorial/opport - The initial value in **Table name** is derived from the row name. You can change it at any time; however, the [Storage bucket](/storage/buckets/) where the table will be saved to cannot be changed. -- **Incremental load** will turn on [incremental loading to Storage](/storage/tables/#incremental-loading). The result of the +- **Incremental load** will turn on [incremental loading to Storage](/storage/tables/incremental-loading/). The result of the incremental load depends on other settings (mainly **Primary Key**). - **Delimiter** and **Enclosure** specify the CSV settings. diff --git a/src/content/docs/components/extractors/storage/storage-api/index.md b/src/content/docs/components/extractors/storage/storage-api/index.md index bb21cdbda..12db17d73 100644 --- a/src/content/docs/components/extractors/storage/storage-api/index.md +++ b/src/content/docs/components/extractors/storage/storage-api/index.md @@ -56,7 +56,7 @@ As the token has access to a single bucket only, you do not need to specify the ### Save Settings -- **Incremental** -- enables [incremental loading](/storage/tables/#incremental-loading) in the current project. If the **Primary Key** is not set, the data is appended. +- **Incremental** -- enables [incremental loading](/storage/tables/incremental-loading/) in the current project. If the **Primary Key** is not set, the data is appended. Otherwise the rows with an existing primary key are updated. - **Primary Key** -- sets the primary key of the table in the current project. The primary key does not have to be the same as in the *source project*. diff --git a/src/content/docs/components/writers/bi-tools/gooddata/index.md b/src/content/docs/components/writers/bi-tools/gooddata/index.md index f42cdc68c..b8cc4039b 100644 --- a/src/content/docs/components/writers/bi-tools/gooddata/index.md +++ b/src/content/docs/components/writers/bi-tools/gooddata/index.md @@ -101,7 +101,7 @@ Incremental load will keep the existing data in the GoodData project. It can be much faster, but the source data needs to be correctly prepared. The incremental load relies on the following two features: -- [Incremental processing](/storage/tables/#incremental-processing) in Storage, and +- [Incremental processing](/storage/tables/incremental-loading/#incremental-processing) in Storage, and - Identity in GoodData; this can be either a `CONNECTION_POINT` column (which acts as a database primary key), or a [Fact Grain](https://help.gooddata.com/doc/enterprise/en/data-integration/data-modeling-in-gooddata/logical-data-model-components-in-gooddata/facts-in-logical-data-models#FactsinLogicalDataModels-FactDatasets) (which acts as a compound unique key). Fact Grain can be used when there is no single identifying column (i.e. there is no `CONNECTION_POINT`) and there is a combination of columns which can be used to identify rows. If there is no such combination, then incremental loading cannot be used. From the table configuration page, you can also **Run** a load of a single table. However, if the table has relations to other diff --git a/src/content/docs/components/writers/bi-tools/tableau/index.md b/src/content/docs/components/writers/bi-tools/tableau/index.md index 1f9b2a527..2865230b7 100644 --- a/src/content/docs/components/writers/bi-tools/tableau/index.md +++ b/src/content/docs/components/writers/bi-tools/tableau/index.md @@ -43,7 +43,7 @@ be part of the TDE file. When configuring the data types, use the *Preview* icon As optional last steps in table configuration, you can configure the name of the TDE file (useful for uploading to Dropbox or Google Drive) and *Table data filter*. The table data filter allows you to set a simple filter for one column or -take advantage of [Incremental processing](/storage/tables/#incremental-processing) by writing only +take advantage of [Incremental processing](/storage/tables/incremental-loading/#incremental-processing) by writing only recently modified data. ![Screenshot - Table Configuration Additional Settings](/components/writers/bi-tools/tableau/tableau-4.png) diff --git a/src/content/docs/components/writers/bi-tools/thoughtspot/index.md b/src/content/docs/components/writers/bi-tools/thoughtspot/index.md index 329c36739..c0dfa73da 100644 --- a/src/content/docs/components/writers/bi-tools/thoughtspot/index.md +++ b/src/content/docs/components/writers/bi-tools/thoughtspot/index.md @@ -58,4 +58,4 @@ In the **Full Load** mode, the table is completely overwritten including the tab using the [`DROP`](https://docs.thoughtspot.com/software/latest/tql-cli-commands) command and it is recreated. Additionally, you can specify a **primary key** for the table, a simple column **data filter**, and a filter for -[incremental processing](/storage/tables/#incremental-processing). +[incremental processing](/storage/tables/incremental-loading/#incremental-processing). diff --git a/src/content/docs/components/writers/database/mssql/index.md b/src/content/docs/components/writers/database/mssql/index.md index 51a01f4e0..381dbf5fb 100644 --- a/src/content/docs/components/writers/database/mssql/index.md +++ b/src/content/docs/components/writers/database/mssql/index.md @@ -69,7 +69,7 @@ This means that if the database is used by other applications which acquire tabl freeze waiting for the locks to be released. Additionally, you can specify a **Primary key** of the table, a simple column **Data filter**, and a filter for -[incremental processing](/storage/tables/#incremental-processing). +[incremental processing](/storage/tables/incremental-loading/#incremental-processing). ## Binary types diff --git a/src/content/docs/components/writers/database/mysql/index.md b/src/content/docs/components/writers/database/mysql/index.md index 2468a84a1..49892a84b 100644 --- a/src/content/docs/components/writers/database/mysql/index.md +++ b/src/content/docs/components/writers/database/mysql/index.md @@ -65,4 +65,4 @@ This means that if the database is used by other applications which acquire tabl freeze waiting for the locks to be released. Additionally, you can specify a **primary key** of the table, a simple column **data filter**, and a filter for -[incremental processing](/storage/tables/#incremental-processing). +[incremental processing](/storage/tables/incremental-loading/#incremental-processing). diff --git a/src/content/docs/components/writers/database/oracle/index.md b/src/content/docs/components/writers/database/oracle/index.md index 4130a0c59..eebed2e98 100644 --- a/src/content/docs/components/writers/database/oracle/index.md +++ b/src/content/docs/components/writers/database/oracle/index.md @@ -67,4 +67,4 @@ fail with the following message: Query failed: 'ORA-00955: name is already used by an existing object Additionally, you can specify a **Primary key** of the table, a simple column **Data filter**, and a filter for -[incremental processing](/storage/tables/#incremental-processing). +[incremental processing](/storage/tables/incremental-loading/#incremental-processing). diff --git a/src/content/docs/components/writers/database/postgresql/index.md b/src/content/docs/components/writers/database/postgresql/index.md index c72eaebba..8d8ababda 100644 --- a/src/content/docs/components/writers/database/postgresql/index.md +++ b/src/content/docs/components/writers/database/postgresql/index.md @@ -68,4 +68,4 @@ freeze waiting for the locks to be released. This will be recorded in the connec Table "account" is locked by 1 transactions, waiting for them to finish Additionally, you can specify a **Primary key** of the table, a simple column **Data filter**, and a filter for -[incremental processing](/storage/tables/#incremental-processing). +[incremental processing](/storage/tables/incremental-loading/#incremental-processing). diff --git a/src/content/docs/components/writers/database/redshift/index.md b/src/content/docs/components/writers/database/redshift/index.md index 277bcb040..ad772676c 100644 --- a/src/content/docs/components/writers/database/redshift/index.md +++ b/src/content/docs/components/writers/database/redshift/index.md @@ -94,7 +94,7 @@ freeze waiting for the locks to be released. See the [Redshift docs](https://doc for more details. Additionally, you can specify a **Primary key** of the table, a simple column **Data filter**, and a filter for -[incremental processing](/storage/tables/#incremental-processing). +[incremental processing](/storage/tables/incremental-loading/#incremental-processing). ## Using Keboola Provisioned Database The connector offers the option to create a [Keboola Provisioned database](#keboola-redshift-database) for you. You can diff --git a/src/content/docs/components/writers/database/snowflake/index.md b/src/content/docs/components/writers/database/snowflake/index.md index a87f50c32..0b3e5814f 100644 --- a/src/content/docs/components/writers/database/snowflake/index.md +++ b/src/content/docs/components/writers/database/snowflake/index.md @@ -95,7 +95,7 @@ using the [`ALTER SWAP`](https://docs.snowflake.com/en/sql-reference/sql/alter-t the shortest unavailability of the target table. However, this operation still drops the table. Additionally, you can specify a **Primary key** of the table, a simple column **Data filter**, and a **Data changed in last** filter for -[incremental processing](/storage/tables/#incremental-processing). +[incremental processing](/storage/tables/incremental-loading/#incremental-processing). **Data changed in last** filter is not available when using **Automatic incremental load**, as the component will use the last run date to determine the data to be loaded. diff --git a/src/content/docs/components/writers/database/synapse/index.md b/src/content/docs/components/writers/database/synapse/index.md index 0857e9924..90e04920f 100644 --- a/src/content/docs/components/writers/database/synapse/index.md +++ b/src/content/docs/components/writers/database/synapse/index.md @@ -90,4 +90,4 @@ This means that if the database is used by other applications which acquire tabl freeze waiting for the locks to be released. Additionally, you can specify a **Primary key** of the table, a simple column **Data filter**, and a filter for -[incremental processing](/storage/tables/#incremental-processing). +[incremental processing](/storage/tables/incremental-loading/#incremental-processing). diff --git a/src/content/docs/components/writers/storage/google-drive/index.md b/src/content/docs/components/writers/storage/google-drive/index.md index 878548300..508eda4a2 100644 --- a/src/content/docs/components/writers/storage/google-drive/index.md +++ b/src/content/docs/components/writers/storage/google-drive/index.md @@ -19,7 +19,7 @@ Then click the **New Table** button to add a new table: ![Screenshot - Add Table Step 1](/components/writers/storage/google-drive/google-drive-1.png) -Select a table from Storage. You may also specify additional filters as well as [incremental processing](/storage/tables/#incremental-processing). +Select a table from Storage. You may also specify additional filters as well as [incremental processing](/storage/tables/incremental-loading/#incremental-processing). All options may be modified later. Click **Next** to select how to load the table to Google Drive: ![Screenshot - Add Table Step 2](/components/writers/storage/google-drive/google-drive-2.png) diff --git a/src/content/docs/components/writers/storage/google-sheets/index.md b/src/content/docs/components/writers/storage/google-sheets/index.md index 10cb525ca..89da98747 100644 --- a/src/content/docs/components/writers/storage/google-sheets/index.md +++ b/src/content/docs/components/writers/storage/google-sheets/index.md @@ -19,7 +19,7 @@ Then click the **New Table** button to add a new table: ![Screenshot - Add Table Step 1](/components/writers/storage/google-sheets/google-sheets-1.png) -Select a table from Storage. You may also specify additional filters as well as [incremental processing](/storage/tables/#incremental-processing). +Select a table from Storage. You may also specify additional filters as well as [incremental processing](/storage/tables/incremental-loading/#incremental-processing). All options may be modified later. Click **Next** to select whether to create a new file or write to an existing one: ![Screenshot - Add Table Step 2](/components/writers/storage/google-sheets/google-sheets-2.png) diff --git a/src/content/docs/components/writers/storage/storage-api/index.md b/src/content/docs/components/writers/storage/storage-api/index.md index 4b3afa3f3..d0da70113 100644 --- a/src/content/docs/components/writers/storage/storage-api/index.md +++ b/src/content/docs/components/writers/storage/storage-api/index.md @@ -49,11 +49,11 @@ Each table has different settings but they are all written to the **same project ### Source - **Table** specifies the table in the *source project*. This value cannot be changed. If you want to write another table, create a new item in the configuration. -- **Changed In Last** allows you to use [incremental processing](/storage/tables/#incremental-processing) to write only the recent part of the data. +- **Changed In Last** allows you to use [incremental processing](/storage/tables/incremental-loading/#incremental-processing) to write only the recent part of the data. ### Destination - **Table Name** -- table name in the *target project* and bucket. -- **Mode** -- One of `Update`, `Replace` or `Recreate`. The update mode enables [incremental loading](/storage/tables/#incremental-loading) +- **Mode** -- One of `Update`, `Replace` or `Recreate`. The update mode enables [incremental loading](/storage/tables/incremental-loading/) in the *target project*. The **primary key** setting will be used from the table in the current project. The replace mode replaces all data in the target table, but keeps the table structure. The Recreate mode drops and creates the target table. Note that the Recreate mode creates a brief moments where the target table does not exist. This can create problems for example when an orchestration diff --git a/src/content/docs/storage/bucket-exposure/index.md b/src/content/docs/storage/bucket-exposure/index.md index 57a38fbc2..0ee5269b8 100644 --- a/src/content/docs/storage/bucket-exposure/index.md +++ b/src/content/docs/storage/bucket-exposure/index.md @@ -1,12 +1,14 @@ --- title: Bucket Exposure slug: 'storage/bucket-exposure' +description: Bucket Exposure shares live Keboola bucket data directly into consumers' Google BigQuery environments — no exports or copies; when to use it and how to set it up. --- :::caution Important: This feature is currently available in BETA and only for BigQuery projects. Contact Keboola support to have it enabled for your project. ::: + diff --git a/src/content/docs/storage/buckets/index.md b/src/content/docs/storage/buckets/index.md index 199988019..e1fe1b41a 100644 --- a/src/content/docs/storage/buckets/index.md +++ b/src/content/docs/storage/buckets/index.md @@ -1,21 +1,22 @@ ---- -title: Buckets -slug: 'storage/buckets' ---- - -Buckets are containers for tables in Storage. They are further organized into the following two **stages**: - -1. **in** — for input data (usually data source connector results) -2. **out** — for processed data (usually results of transformations or applications) - -The distinction between the input and output stages is purely conventional differentiation between raw and processed data. -When creating a new bucket, select one of the stages and a suitable [database backend](/storage/#storage-backend-types-and-features) based on its properties. -For information on how to load data into Storage, see the corresponding part of our [tutorial](/tutorial/load/). - -![Screenshot - Create bucket](/storage/buckets/create-bucket.png) - -To review information about an existing bucket, hover over the bucket name and select **Bucket detail**: - -![Screenshot - Bucket information](/storage/buckets/bucket-info.png) - -Apart from being used for organizing tables, buckets can also be used for [sharing tables](/catalog/). +--- +title: Buckets +slug: 'storage/buckets' +description: Buckets are containers for tables in Storage, organized into in and out stages — creating a bucket, reviewing its detail, and using buckets to share tables. +--- + +Buckets are containers for tables in Storage. They are further organized into the following two **stages**: + +1. **in** — for input data (usually data source connector results) +2. **out** — for processed data (usually results of transformations or applications) + +The distinction between the input and output stages is purely conventional differentiation between raw and processed data. +When creating a new bucket, select one of the stages and a suitable [database backend](/storage/#storage-backend-types-and-features) based on its properties. +For information on how to load data into Storage, see the corresponding part of our [tutorial](/tutorial/load/). + +![Screenshot - Create bucket](/storage/buckets/create-bucket.png) + +To review information about an existing bucket, click the bucket name in the **Tables & Buckets** tab; the bucket detail shows its ID, backend, sharing status, schema, stage, and size: + +![Screenshot - Bucket information](/storage/buckets/bucket-info.png) + +Apart from being used for organizing tables, buckets can also be used for [sharing tables](/catalog/). diff --git a/src/content/docs/storage/byobq/byobq.md b/src/content/docs/storage/byobq/byobq.md index 4ccb98343..bd2b73b65 100644 --- a/src/content/docs/storage/byobq/byobq.md +++ b/src/content/docs/storage/byobq/byobq.md @@ -1,6 +1,7 @@ --- title: How to Connect BigQuery slug: 'storage/byobq' +description: Set up Google Cloud resources for a BigQuery backend — create the folder, project, service account, and permissions Keboola needs to connect BigQuery to your project. --- diff --git a/src/content/docs/storage/byodb.md b/src/content/docs/storage/byodb.md index 1a87cd533..8ebf19f9a 100644 --- a/src/content/docs/storage/byodb.md +++ b/src/content/docs/storage/byodb.md @@ -1,6 +1,7 @@ --- title: Bring Your Own Database (BYODB) slug: 'storage/byodb' +description: 'Bring Your Own Database — host Keboola project data in your own Snowflake or BigQuery account: access roles, dynamic backend sizes, and rules for Keboola-created objects.' --- In some instances, you can use your own Snowflake/BigQuery account to host data from Keboola. Currently, this option is only supported for the Snowflake and BigQuery backends. diff --git a/src/content/docs/storage/byodb/external-buckets/index.md b/src/content/docs/storage/byodb/external-buckets/index.md index 695f9f6bb..6b60e9dbd 100644 --- a/src/content/docs/storage/byodb/external-buckets/index.md +++ b/src/content/docs/storage/byodb/external-buckets/index.md @@ -1,6 +1,7 @@ --- title: External Datasets slug: 'storage/byodb/external-buckets' +description: External Datasets register an existing Snowflake schema or BigQuery dataset as a read-only virtual bucket in Storage — requirements, registration steps, and limits. --- diff --git a/src/content/docs/storage/byodb/snowflake-secure-data-sharing/index.md b/src/content/docs/storage/byodb/snowflake-secure-data-sharing/index.md index 42647d13b..ec13fc877 100644 --- a/src/content/docs/storage/byodb/snowflake-secure-data-sharing/index.md +++ b/src/content/docs/storage/byodb/snowflake-secure-data-sharing/index.md @@ -1,6 +1,7 @@ --- title: Snowflake Secure Data Sharing slug: 'storage/byodb/snowflake-secure-data-sharing' +description: Share Snowflake data with a Keboola BYODB account via Snowflake Secure Data Sharing — producer and consumer workflows for creating and consuming shares. --- :::caution diff --git a/src/content/docs/storage/data-streams/data-streams.md b/src/content/docs/storage/data-streams/data-streams.md index b8ab84783..f4c2b17fe 100644 --- a/src/content/docs/storage/data-streams/data-streams.md +++ b/src/content/docs/storage/data-streams/data-streams.md @@ -4,6 +4,7 @@ slug: 'storage/data-streams' redirect_from: - /integrate/data-streams/ - /integrate/push-data/ +description: Data Streams push event data straight into Storage over HTTP or OpenTelemetry — no connector or middleware needed; sources, import conditions, and setup. --- @@ -66,7 +67,7 @@ table and no deduplication happens on import. If you need unique rows (e.g., one row per event ID), it is up to you to decide whether deduplication is required and to handle it **downstream**: create a -[deduplication transformation](/storage/tables/#primary-key-deduplication) that reads the stream +[deduplication transformation](/storage/tables/incremental-loading/#primary-key-deduplication) that reads the stream table and writes the result to a new table where the [output mapping](/transformations/mappings/#output-mapping) primary key performs the deduplication on load. Schedule it to run regularly with a [conditional flow](/flows/). diff --git a/src/content/docs/storage/data-streams/opentelemetry/index.md b/src/content/docs/storage/data-streams/opentelemetry/index.md index 90b218de1..cd1b78e18 100644 --- a/src/content/docs/storage/data-streams/opentelemetry/index.md +++ b/src/content/docs/storage/data-streams/opentelemetry/index.md @@ -1,6 +1,7 @@ --- title: OpenTelemetry (OTLP) Data Streams slug: 'storage/data-streams/opentelemetry' +description: Use the OpenTelemetry (OTLP) source type to send logs, metrics, and traces from any OTel SDK or collector into Keboola Storage and query telemetry beside business data. --- diff --git a/src/content/docs/storage/files/index.md b/src/content/docs/storage/files/index.md index 221f20c40..1cca5b440 100644 --- a/src/content/docs/storage/files/index.md +++ b/src/content/docs/storage/files/index.md @@ -1,13 +1,14 @@ --- title: Files slug: 'storage/files' +description: File Storage in Keboola — uploading files with tags and retention options, public and non-public download links, sliced files, and file size limits. redirect_from: - /storage/file-uploads/ --- -The *File Storage* is available in the **Files** section of Storage and contains all raw files uploaded to your project. +The *File Storage* is available in the **Files** tab of Storage and contains all raw files uploaded to your project. It also contains files with data exported from tables. These are created when you request to export a table from *Table Storage*. The *File Storage* serves two main purposes: 1. Files can be used to store an arbitrary file. @@ -69,6 +70,7 @@ This can happen for some [exported or imported tables](/storage/tables/uploads/) Merging a sliced file requires a [substantial effort](https://developers.keboola.com/integrate/storage/api/import-export/#working-with-sliced-files). ## Limits + The maximum allowed size of an uploaded file is currently 2 GB (2,048,000,000 bytes exactly). This applies to both file and table uploads. The actual table size may be bigger, because the table is uploaded as a compressed file. diff --git a/src/content/docs/storage/index.md b/src/content/docs/storage/index.md index c0fe8f5d7..b3a73a011 100644 --- a/src/content/docs/storage/index.md +++ b/src/content/docs/storage/index.md @@ -1,64 +1,65 @@ --- title: Storage slug: 'storage' ---- - - - -*See our [Getting Started](/tutorial/load/) tutorial for instructions on how to use Storage.* - -As the central [Keboola subsystem](/overview/), Storage manages everything related to **storing** data and **accessing** it. -It is implemented as a layer on top of database engines that we use as our backends -([Snowflake](https://www.snowflake.com/), [BigQuery](https://cloud.google.com/bigquery)). -By default all new [Pay As You Go projects](/management/payg-project/) use the BigQuery backend. If you are a contract customer, you can select which backend you want to use. - -As with all other Keboola components, everything that can be done through the UI can be also done programmatically -via the [Storage API](https://api.keboola.com/?service=storage). -See our [developers guide](https://developers.keboola.com/integrate/storage/) to learn more. -Every Storage operation must be authorized via a [token](/management/project/tokens/). -It is also recorded in [Events](/management/project/tokens/#token-events) and -[Jobs](/management/jobs/). - -## Storage Data -The Storage component manages all data stored in each Keboola project: - -- [Data tables](/storage/tables/) (Table Storage) — organized into [buckets](/storage/buckets/) -- [Data files](/storage/files/) (File Storage) — all raw files uploaded to your project -- [Component configurations](/components/) - -Different storage technologies are used for the above data: - -- [Amazon S3 Storage](https://aws.amazon.com/s3/), [Azure Blob Storage](https://azure.microsoft.com/en-us/services/storage/blobs/), and [Google Cloud Storage](https://cloud.google.com/storage/) are used for [File Storage](/storage/files/). -- [BigQuery](https://cloud.google.com/bigquery/) and [Snowflake](https://www.snowflake.com/product/) are used for [Table Storage](/storage/tables/). - -The database system behind the Table Storage is referred to as a **backend**. -Data in Table Storage is internally stored in a **database backend** (project backend). - -### Storage Backend Types and Features -There are multiple types of storage backend configurations you can choose from, and each has its own features and limitations. - -#### Multi-tenant (MT) -MT backend shares compute resources across projects, while maintaining a high degree of data isolation between those projects. Benefits of this setup are mainly cost-savings, as projects share the compute resources. - -**Snowflake** projects are *not allowed* to access the Snowflake database directly through [Snowflake Workspaces](/workspace/#snowflake). This limitation exists because all databases sharing the same resources exist in the same Snowflake account, and Snowflake does not allow direct access to the same account by different business entities. Instead, users can use the in-platform SQL Editor for SQL development and analysis above the data in Table Storage. The Data Gateway component can also be used to share data in read-only mode with third-party BI and visualization tools. - -**BigQuery** projects are *allowed* to access BigQuery database directly through [BigQuery Workspaces](/workspace/#bigquery), as BigQuery handles access to its environment differently. - -MT backend type is used for all [Pay As You Go projects](/management/payg-project/) and most contract customers who don't need extra features coming with other backend types. - -#### Keboola-Brings-Database (KBDB) -KBDB is **Snowflake** backend type where Keboola manages a dedicated account for a customer, while the customer is not an admin of the account. This setup is used for contract customers who need access to [Snowflake Workspaces](/workspace/#snowflake) and direct connection to Snowflake database (e.g. for SQL development in IDEs). - -As the dedicated account does not share the resources with other customers, it has higher and more stable performance, but also higher cost. - - -#### Bring-Your-Own-Database (BYODB) -This backend type allows customers to use their own database for their Keboola projects and can use both BigQuery and Snowflake as the backend. - -As the database backend is owned by the customer, there are no functional limitations and it's also fully managed and paid for by the customer. - -To learn more about the BYODB setup, please refer to the [BigQuery](/storage/byobq/) and [Snowflake](/storage/byodb/#snowflake) BYODB documentation. - ---- - +description: Keboola Storage — tables, files, and configurations on Snowflake or BigQuery backends, and the multi-tenant, Keboola-managed, and bring-your-own-database backend types. +--- + + + +*See our [Getting Started](/tutorial/load/) tutorial for instructions on how to use Storage.* + +As the central [Keboola subsystem](/overview/), Storage manages everything related to **storing** data and **accessing** it. +It is implemented as a layer on top of database engines that we use as our backends +([Snowflake](https://www.snowflake.com/), [BigQuery](https://cloud.google.com/bigquery)). +By default all new [Pay As You Go projects](/management/payg-project/) use the BigQuery backend. If you are a contract customer, you can select which backend you want to use. + +As with all other Keboola components, everything that can be done through the UI can be also done programmatically +via the [Storage API](https://api.keboola.com/?service=storage). +See our [developers guide](https://developers.keboola.com/integrate/storage/) to learn more. +Every Storage operation must be authorized via a [token](/management/project/tokens/). +It is also recorded in [Events](/management/project/tokens/#token-events) and +[Jobs](/management/jobs/). + +## Storage Data +The Storage component manages all data stored in each Keboola project: + +- [Data tables](/storage/tables/) (Table Storage) — organized into [buckets](/storage/buckets/) +- [Data files](/storage/files/) (File Storage) — all raw files uploaded to your project +- [Component configurations](/components/) + +Different storage technologies are used for the above data: + +- [Amazon S3 Storage](https://aws.amazon.com/s3/), [Azure Blob Storage](https://azure.microsoft.com/en-us/services/storage/blobs/), and [Google Cloud Storage](https://cloud.google.com/storage/) are used for [File Storage](/storage/files/). +- [BigQuery](https://cloud.google.com/bigquery/) and [Snowflake](https://www.snowflake.com/product/) are used for [Table Storage](/storage/tables/). + +The database system behind the Table Storage is referred to as a **backend**. +Data in Table Storage is internally stored in a **database backend** (project backend). + +### Storage Backend Types and Features +There are multiple types of storage backend configurations you can choose from, and each has its own features and limitations. + +#### Multi-tenant (MT) +MT backend shares compute resources across projects, while maintaining a high degree of data isolation between those projects. Benefits of this setup are mainly cost-savings, as projects share the compute resources. + +**Snowflake** projects are *not allowed* to access the Snowflake database directly through [Snowflake Workspaces](/workspace/#snowflake). This limitation exists because all databases sharing the same resources exist in the same Snowflake account, and Snowflake does not allow direct access to the same account by different business entities. Instead, users can use the in-platform SQL Editor for SQL development and analysis above the data in Table Storage. The Data Gateway component can also be used to share data in read-only mode with third-party BI and visualization tools. + +**BigQuery** projects are *allowed* to access BigQuery database directly through [BigQuery Workspaces](/workspace/#bigquery), as BigQuery handles access to its environment differently. + +MT backend type is used for all [Pay As You Go projects](/management/payg-project/) and most contract customers who don't need extra features coming with other backend types. + +#### Keboola-Brings-Database (KBDB) +KBDB is **Snowflake** backend type where Keboola manages a dedicated account for a customer, while the customer is not an admin of the account. This setup is used for contract customers who need access to [Snowflake Workspaces](/workspace/#snowflake) and direct connection to Snowflake database (e.g. for SQL development in IDEs). + +As the dedicated account does not share the resources with other customers, it has higher and more stable performance, but also higher cost. + + +#### Bring-Your-Own-Database (BYODB) +This backend type allows customers to use their own database for their Keboola projects and can use both BigQuery and Snowflake as the backend. + +As the database backend is owned by the customer, there are no functional limitations and it's also fully managed and paid for by the customer. + +To learn more about the BYODB setup, please refer to the [BigQuery](/storage/byobq/) and [Snowflake](/storage/byodb/#snowflake) BYODB documentation. + +--- + If you have any questions regarding the backend types or need help with the setup, please contact our [support](/management/support/). \ No newline at end of file diff --git a/src/content/docs/storage/jobs/index.md b/src/content/docs/storage/jobs/index.md index 2ee63eb5b..5e9c295dd 100644 --- a/src/content/docs/storage/jobs/index.md +++ b/src/content/docs/storage/jobs/index.md @@ -7,14 +7,14 @@ description: The low-level record of data loaded to and unloaded from Table Stor Storage jobs are not to be confused with [component jobs](/management/jobs/). Storage jobs represent a low level view of -all data loaded to and unloaded from the [Table Storage](/storage/tables). Working with the Storage jobs is rarely necessary +all data loaded to and unloaded from the [Table Storage](/storage/tables/). Working with the Storage jobs is rarely necessary for end-users. We provide this view mainly to maintain complete transparency of what is happening in your project. ![Screenshot - Create alias](/storage/jobs/storage-jobs-1.png) The view of Storage jobs can be useful in very busy projects, or in cases where you see a component [job](/management/jobs/) being stuck on the message `Waiting for X Storage jobs to finish.`. There is a core limitation of Storage Tables — only one job may -write to a table at a time. So when you see jobs taking longer than usual or waiting for Storage Jobs to finish, you might want to +write to a table at a time. So when you see jobs taking longer than usual or waiting for Storage Jobs to finish, you might want to check the Storage Jobs to understand what is happening. You can see how many jobs are processing if they are writing to same tables, and then, for example, adjust flow triggers to avoid concurrency issues. diff --git a/src/content/docs/storage/tables/backups.md b/src/content/docs/storage/tables/backups.md index 1ac672f22..52f78d44a 100644 --- a/src/content/docs/storage/tables/backups.md +++ b/src/content/docs/storage/tables/backups.md @@ -1,6 +1,7 @@ --- title: Backups and Restorations slug: 'storage/tables/backups' +description: Back up and restore Storage tables — time-travel restore on the Snowflake backend and table snapshots on all backends. --- @@ -9,6 +10,7 @@ There are two methods available for backing up and restoring data: 1. Time travel restore – available only for the Snowflake backend 2. Snapshots – available for all backends + ![Screenshot - Storage Backups](/storage/tables/snap-restore.png) diff --git a/src/content/docs/storage/tables/csv-files.md b/src/content/docs/storage/tables/csv-files.md index 532a744a5..b9d36887d 100644 --- a/src/content/docs/storage/tables/csv-files.md +++ b/src/content/docs/storage/tables/csv-files.md @@ -1,6 +1,7 @@ --- title: CSV Files slug: 'storage/tables/csv-files' +description: The CSV formats Storage accepts and produces — delimiters, enclosures, encoding, line breaks, and step-by-step CSV export/import with Microsoft Excel. --- diff --git a/src/content/docs/storage/tables/data-types/index.md b/src/content/docs/storage/tables/data-types/index.md index 2b475f69b..807b4d716 100644 --- a/src/content/docs/storage/tables/data-types/index.md +++ b/src/content/docs/storage/tables/data-types/index.md @@ -1,207 +1,208 @@ --- title: Native Data Types slug: 'storage/tables/data-types' ---- - - - -The **Native Data Types** feature streamlines the process of **propagating data types** from the source to the storage. With Keboola Native Data Types, the system automatically maintains data types throughout the pipeline, eliminating the need for manual intervention and reducing errors in data processing. - -For example, when a user imports a large dataset with predefined data types, such as NUMERIC, BOOLEAN, and DATE, these types are preserved automatically. Without Native Data Types, the data would have been imported as VARCHAR, requiring the user to manually update the types in transformation. - -Non-typed tables without native data types are labeled in the UI with a badge: **non-typed**. - -## Key Benefits -These are the key benefits of using the Native Data Types feature: -- **Automatic Data Type Preservation:** Data types from the source are automatically respected, reducing the need for manual adjustments in Storage. -- **Faster Data Handling:** Native data types enable more efficient data manipulation, as well as faster loading and unloading, improving overall performance. -- **Simplified Transformations:** Read-only data access eliminates the need for casting, making data operations smoother and more streamlined. -- **Flexible Configurations:** Users can decide whether data types should be automatically fetched for each configuration when creating a table. -- **Improved Workspace Loading:** Loading data into a workspace is significantly faster than loading into a table without native data types, eliminating the need for additional casting. -- **Typed Columns in Workspaces:** Tables **accessed in a workspace** via the [read-only input mapping](/workspace/#read-only-input-mapping) already have typed columns, ensuring seamless data handling. - -## Current Drawbacks -Using the Native Data Types feature also has its drawbacks: -- Data types in typed tables cannot be modified after creation. To change the data types, you must recreate the table. This limitation applies to both the UI and the API. See [How to Change Column Types](/storage/tables/data-types/#changing-types-of-existing-typed-columns). -- Keboola does not perform any type conversion during data loading. Your data must exactly match the column type defined in the table within Storage. -- Loading data with incompatible types will result in a failure. - -## How It Works -By default, all new tables are created as typed tables if the component supports this feature. Non-typed tables are labeled in the Storage UI with the label NON-TYPED. - -You can configure the data type behavior in the UI component configuration settings. If the component supports this feature, you will see the option **Automatic data types** in the right menu, which can be toggled ON and OFF. -- **When enabled:** The component creates a typed table that respects the data types from the source (e.g., DATETIME, BOOLEAN). -- **When disabled:** A typed table is created with all columns as VARCHAR, and data types are stored as metadata. - -In transformations, this option is not available. Instead, you define the data types in your query (if you need the table to be typed). If no types are defined, the table will default to storing data in VARCHAR format. - -**Important:** Existing tables will not be affected by this feature. Also, if you do not see the **Automatic data types** option in the sidebar, it means the component does not support this feature. - -### How to Create a Typed Table -The Native Data Types feature allows tables to be created with data types that match the original source or storage backend. Here’s how you can create typed tables: -- **Manually via API** -You can manually create typed tables using the [tables-definition endpoint](https://keboola.docs.apiary.io/#reference/tables/create-table-definition/create-new-table-definition). Ensure that the data types align with the storage backend (e.g., Snowflake, BigQuery) used in your project. Alternatively, [base types](/storage/tables/data-types/#base-types) can be used for compatibility. -- **Using a Component** -Extractors and transformations that match the storage backend (e.g., Snowflake SQL transformation on a Snowflake storage backend) will automatically create typed tables in Storage: - - **Matching Storage Backend:** Database extractors and transformations create storage tables using the same data types as the backend. - - **Mismatching Storage Backend:** Extractors use base types to ensure compatibility. [Learn more.](/storage/tables/data-types/#base-types) - - +description: Native Data Types preserve source column types (NUMERIC, DATE, BOOLEAN, …) end-to-end in Storage — benefits, drawbacks, and how to change types on typed tables. +--- + + + +The **Native Data Types** feature streamlines the process of **propagating data types** from the source to the storage. With Keboola Native Data Types, the system automatically maintains data types throughout the pipeline, eliminating the need for manual intervention and reducing errors in data processing. + +For example, when a user imports a large dataset with predefined data types, such as NUMERIC, BOOLEAN, and DATE, these types are preserved automatically. Without Native Data Types, the data would have been imported as VARCHAR, requiring the user to manually update the types in transformation. + +Non-typed tables without native data types are labeled in the UI with a badge: **non-typed**. + +## Key Benefits +These are the key benefits of using the Native Data Types feature: +- **Automatic Data Type Preservation:** Data types from the source are automatically respected, reducing the need for manual adjustments in Storage. +- **Faster Data Handling:** Native data types enable more efficient data manipulation, as well as faster loading and unloading, improving overall performance. +- **Simplified Transformations:** Read-only data access eliminates the need for casting, making data operations smoother and more streamlined. +- **Flexible Configurations:** Users can decide whether data types should be automatically fetched for each configuration when creating a table. +- **Improved Workspace Loading:** Loading data into a workspace is significantly faster than loading into a table without native data types, eliminating the need for additional casting. +- **Typed Columns in Workspaces:** Tables **accessed in a workspace** via the [read-only input mapping](/workspace/#read-only-input-mapping) already have typed columns, ensuring seamless data handling. + +## Current Drawbacks +Using the Native Data Types feature also has its drawbacks: +- Data types in typed tables cannot be modified after creation. To change the data types, you must recreate the table. This limitation applies to both the UI and the API. See [How to Change Column Types](/storage/tables/data-types/#changing-types-of-existing-typed-columns). +- Keboola does not perform any type conversion during data loading. Your data must exactly match the column type defined in the table within Storage. +- Loading data with incompatible types will result in a failure. + +## How It Works +By default, all new tables are created as typed tables if the component supports this feature. Non-typed tables are labeled in the Storage UI with the label NON-TYPED. + +You can configure the data type behavior in the UI component configuration settings. If the component supports this feature, you will see the option **Automatic data types** in the right menu, which can be toggled ON and OFF. +- **When enabled:** The component creates a typed table that respects the data types from the source (e.g., DATETIME, BOOLEAN). +- **When disabled:** A typed table is created with all columns as VARCHAR, and data types are stored as metadata. + +In transformations, this option is not available. Instead, you define the data types in your query (if you need the table to be typed). If no types are defined, the table will default to storing data in VARCHAR format. + +**Important:** Existing tables will not be affected by this feature. Also, if you do not see the **Automatic data types** option in the sidebar, it means the component does not support this feature. + +### How to Create a Typed Table +The Native Data Types feature allows tables to be created with data types that match the original source or storage backend. Here’s how you can create typed tables: +- **Manually via API** +You can manually create typed tables using the [tables-definition endpoint](https://keboola.docs.apiary.io/#reference/tables/create-table-definition/create-new-table-definition). Ensure that the data types align with the storage backend (e.g., Snowflake, BigQuery) used in your project. Alternatively, [base types](/storage/tables/data-types/#base-types) can be used for compatibility. +- **Using a Component** +Extractors and transformations that match the storage backend (e.g., Snowflake SQL transformation on a Snowflake storage backend) will automatically create typed tables in Storage: + - **Matching Storage Backend:** Database extractors and transformations create storage tables using the same data types as the backend. + - **Mismatching Storage Backend:** Extractors use base types to ensure compatibility. [Learn more.](/storage/tables/data-types/#base-types) + + :::caution **Important:** When a table is created, it defaults to the lengths and precisions specific to the Storage backend. For instance, in Snowflake, the NUMBER base type defaults to NUMBER(38,9), which might differ from the source database column type, such as NUMBER(10,2). To avoid this limitation, follow the steps below. -::: - -To avoid the limitation: -- Manually create the table in advance using the [Table Definition API](https://keboola.docs.apiary.io/#reference/tables/create-table-definition/create-new-table-definition), specifying the correct lengths and precisions. -- Subsequent jobs writing data to this table will respect your defined schema as long as it matches the expected structure. -- Be cautious when dropping and recreating tables. If a job creates a table, it will default to the base type with backend-specific defaults, which might not align with your source. - -**Example:** -To ensure typed tables are imported correctly into Storage, define your table in a Snowflake SQL transformation, adhering to the desired schema and data types: - -![Screenshot - Create a Table](/storage/tables/data-types/create-table.png) - -## Base Types -Source data types are mapped to a destination using a **base type**. The current base types are `STRING`, `INTEGER`, `NUMERIC`, `FLOAT`, `BOOLEAN`, `DATE`, and `TIMESTAMP`. For example, a MySQL extractor may store a column with the data type `BIGINT`. This type is mapped to the `INTEGER` base type, ensuring high interoperability between components. - -For detailed mappings, please refer to the [conversion table](https://developers.keboola.com/extend/common-interface/manifest-files/out-tables-manifests-native-types/#data-type-conversions). You can also view the extracted data types in the [storage table](/storage/tables/) detail. - -### How to Define Data Types - -#### Using actual data types of the storage backend -For example, in the case of Snowflake, you can create a column with a specific type like `TIMESTAMP_NTZ` or `DECIMAL(20,2)`. This approach allows you to define all details of the data type, including precision and scale. An example of such a column definition in a table-definition API endpoint call might look like this: - -``` -{ - "name": "id", - "definition": { - "type": "DECIMAL", - "length": "20,2", - "nullable": false, - "default": "999" - } -} -``` - -#### Using Keboola-provided base types -Specifying native types using Keboola’s [base types](/storage/tables/data-types/#base-types) is ideal for component-provided types, as these are storage backend agnostic. This method ensures compatibility across different storage backends. Additionally, base types can also be used when defining tables via the table-definition API endpoint. The definition format is as follows: - -``` -{ - "name": "id", - "basetype": "NUMERIC" -} -``` - -### Changing Types of Existing Typed Columns -You **cannot change the type of a column in a typed table once it has been created**. However, there are multiple workarounds to address this limitation: - -**For tables using full load:** Drop the table and create a new one with the correct types. Then, load the data into the newly created table. - -**For tables loaded incrementally:** You will need to create a new column with the desired type and migrate the data step by step: - - Assume you have a column `date` of type `VARCHAR` in a typed table, and you want to change it to `TIMESTAMP`. - - Start by adding a new column named `date_timestamp` of type `TIMESTAMP` to the table. - - Update all jobs filling the table to populate both the new column (`date_timestamp`) and the existing column (`date`). - - Run an ad-hoc transformation to copy data from `date` to `date_timestamp` for the existing rows. - - Gradually update all configurations and references to use `date_timestamp` instead of `date`. - - Once all references are updated and the old column is no longer in use, you can safely remove the `date` column. - - +::: + +To avoid the limitation: +- Manually create the table in advance using the [Table Definition API](https://keboola.docs.apiary.io/#reference/tables/create-table-definition/create-new-table-definition), specifying the correct lengths and precisions. +- Subsequent jobs writing data to this table will respect your defined schema as long as it matches the expected structure. +- Be cautious when dropping and recreating tables. If a job creates a table, it will default to the base type with backend-specific defaults, which might not align with your source. + +**Example:** +To ensure typed tables are imported correctly into Storage, define your table in a Snowflake SQL transformation, adhering to the desired schema and data types: + +![Screenshot - Create a Table](/storage/tables/data-types/create-table.png) + +## Base Types +Source data types are mapped to a destination using a **base type**. The current base types are `STRING`, `INTEGER`, `NUMERIC`, `FLOAT`, `BOOLEAN`, `DATE`, and `TIMESTAMP`. For example, a MySQL extractor may store a column with the data type `BIGINT`. This type is mapped to the `INTEGER` base type, ensuring high interoperability between components. + +For detailed mappings, please refer to the [conversion table](https://developers.keboola.com/extend/common-interface/manifest-files/out-tables-manifests-native-types/#data-type-conversions). You can also view the extracted data types in the [storage table](/storage/tables/) detail. + +### How to Define Data Types + +#### Using actual data types of the storage backend +For example, in the case of Snowflake, you can create a column with a specific type like `TIMESTAMP_NTZ` or `DECIMAL(20,2)`. This approach allows you to define all details of the data type, including precision and scale. An example of such a column definition in a table-definition API endpoint call might look like this: + +``` +{ + "name": "id", + "definition": { + "type": "DECIMAL", + "length": "20,2", + "nullable": false, + "default": "999" + } +} +``` + +#### Using Keboola-provided base types +Specifying native types using Keboola’s [base types](/storage/tables/data-types/#base-types) is ideal for component-provided types, as these are storage backend agnostic. This method ensures compatibility across different storage backends. Additionally, base types can also be used when defining tables via the table-definition API endpoint. The definition format is as follows: + +``` +{ + "name": "id", + "basetype": "NUMERIC" +} +``` + +### Changing Types of Existing Typed Columns +You **cannot change the type of a column in a typed table once it has been created**. However, there are multiple workarounds to address this limitation: + +**For tables using full load:** Drop the table and create a new one with the correct types. Then, load the data into the newly created table. + +**For tables loaded incrementally:** You will need to create a new column with the desired type and migrate the data step by step: + - Assume you have a column `date` of type `VARCHAR` in a typed table, and you want to change it to `TIMESTAMP`. + - Start by adding a new column named `date_timestamp` of type `TIMESTAMP` to the table. + - Update all jobs filling the table to populate both the new column (`date_timestamp`) and the existing column (`date`). + - Run an ad-hoc transformation to copy data from `date` to `date_timestamp` for the existing rows. + - Gradually update all configurations and references to use `date_timestamp` instead of `date`. + - Once all references are updated and the old column is no longer in use, you can safely remove the `date` column. + + :::caution **Important:** Always verify other configurations that depend on the table to avoid schema mismatches. Also, pay special attention to writers (data destination connectors), particularly if the table already exists in the destination system. Mismatched schemas between the source and destination can lead to errors. -::: - -### How to Create a Typed Table Based on a Non-Typed Table -If you have a non-typed table, `non_typed_table`, with undefined data types and want to convert it into a typed table, follow these steps: - -**Step 1: Set Up the Transformation** -- Create a new transformation in Keboola. -- Choose `non_typed_table` as the input table in the input mapping section (you can also rely on [read-only input mapping](/transformations/#read-only-input-mapping)). -- In the output mapping section, define the output table as `typed_table`. Ensure that the output table does not exist; otherwise, it will not be created as a typed table. - -![Screenshot - Typed Table Transformation](/storage/tables/data-types/typed-table-transformation.png) - -**Step 2: Define the Query** -In the queries section, write an SQL query to transform the column types. Use proper casting for each column to match the desired data types. - -For example, if you need to format a date column, include the appropriate SQL casting or formatting function in your query. - -``` -CREATE TABLE "typed_table" AS - SELECT - CAST(ntt."id" AS VARCHAR(64)) AS "id", - CAST(ntt."id_profile" AS INTEGER) AS "id_profile", - TO_TIMESTAMP(ntt."date", 'DD.MM.YYYY"T"HH24:MI:SS') AS "date", - CAST(ntt."amount" AS INTEGER) AS "amount" - FROM "non_typed_table" AS ntt; -``` - -**Step 3: Run the Transformation** -Execute the transformation and wait for it to complete. - -**Step 4: Verify the Schema** -Once the transformation is finished, check the schema of the newly created table, `typed_table`. It should now include the appropriate data types. - -***Note:** [Incremental loading](/storage/tables/#incremental-loading) cannot be used when creating a typed table in this manner.* - -## Incremental Loading -The behavior of incremental loading differs between **typed** and **non-typed tables**: - -- **Typed tables:** Only the columns in the table's **primary key** are compared to detect changes. -- **Non-typed tables:** The entire row is compared, and rows are updated if **any value** has changed. - -For more information, refer to our documentation on [incremental loading](/storage/tables/#difference-between-tables-with-native-datatypes-and-string-tables). - -## Handling NULLs -Data can contain `NULL` values or empty strings, which are converted differently based on the processing backend. - -### CSV Format: Unquoted vs Quoted Empty Values -In CSV files, there are two ways to represent an empty value: - -- **Unquoted empty** (`,,`) — the field is completely absent between delimiters -- **Quoted empty** (`""`) — the field explicitly contains an empty string - -This distinction matters because backends treat them differently. - -### Backend Behavior - -| CSV value | Snowflake (STRING) | BigQuery (STRING) | BigQuery (INT64/NUMERIC) | -|---|---|---|---| -| `,,` (unquoted empty) | `NULL` | `NULL` | `NULL` | -| `""` (quoted empty) | `NULL` | `’’` (empty string) | `NULL` | - -**Key difference:** Snowflake converts both `,,` and `""` to `NULL` for string columns. BigQuery preserves `""` as an empty string for STRING columns but coerces it to `NULL` for numeric columns (since an empty string is not a valid number). - -This means the **same CSV file** can produce different NULL semantics depending on the target column type in BigQuery. Numeric columns will have `NULL` where string columns will have an empty string, even though the source CSV value is the same. - -**Note:** This is a behavior of BigQuery's CSV parser, not Keboola Storage. BigQuery does not provide any option (similar to Snowflake's `EMPTY_FIELD_AS_NULL`) to change how quoted empty strings are handled during CSV loading. There is no available workaround at the BigQuery level. - -### The `treatValuesAsNull` Parameter -When importing data, you can use the `treatValuesAsNull` parameter to specify which values should be treated as `NULL`. For example, `treatValuesAsNull: [""]` tells the system to treat empty strings as `NULL`. - -However, in BigQuery, this parameter maps to BigQuery’s native [`nullMarker`](https://cloud.google.com/bigquery/docs/loading-data-cloud-storage-csv#csv-options) CSV load option, which **only applies to unquoted CSV fields**. Quoted values bypass the null marker check entirely — this is a limitation of BigQuery’s CSV parser, not Keboola Storage. - -In practice: - -| CSV value | `treatValuesAsNull: [""]` | BigQuery STRING result | -|---|---|---| -| `,,` (unquoted empty) | matches null marker → **NULL** | `NULL` | -| `""` (quoted empty) | skipped (quoted) | `’’` (empty string) | - -If your data source (e.g., a database extractor) writes `NULL` values as `""` instead of `,,`, string columns will contain empty strings instead of `NULL` in BigQuery, while numeric columns will still show `NULL` due to type coercion. To ensure correct NULL handling, the data source must represent NULL values as unquoted empty fields (`,,`) in the CSV. - -### NULL and Incremental Loading -Columns without native types are always `VARCHAR NOT NULL`. This means you don’t need to worry about specific `NULL` behavior. However, this changes with typed columns. - -In most databases, `NULL` does not equal `NULL` (`NULL == NULL` is not `TRUE`, but `NULL`). This behavior can disrupt the incremental loading process, where columns are compared to detect changes. - -To avoid such issues, ensure that your primary key columns are **not nullable**. This is especially relevant in `CTAS` (Create Table As Select) queries, where columns are nullable by default. To address this, explicitly define the columns as non-nullable in the `CTAS` expression. For example: - -``` -CREATE TABLE "ctas_table" ( - "id" NUMBER NOT NULL, - "name" VARCHAR(255) NOT NULL, - "created_at" TIMESTAMP_NTZ NOT NULL -) AS SELECT * FROM "typed_table"; -``` - +::: + +### How to Create a Typed Table Based on a Non-Typed Table +If you have a non-typed table, `non_typed_table`, with undefined data types and want to convert it into a typed table, follow these steps: + +**Step 1: Set Up the Transformation** +- Create a new transformation in Keboola. +- Choose `non_typed_table` as the input table in the input mapping section (you can also rely on [read-only input mapping](/transformations/#read-only-input-mapping)). +- In the output mapping section, define the output table as `typed_table`. Ensure that the output table does not exist; otherwise, it will not be created as a typed table. + +![Screenshot - Typed Table Transformation](/storage/tables/data-types/typed-table-transformation.png) + +**Step 2: Define the Query** +In the queries section, write an SQL query to transform the column types. Use proper casting for each column to match the desired data types. + +For example, if you need to format a date column, include the appropriate SQL casting or formatting function in your query. + +``` +CREATE TABLE "typed_table" AS + SELECT + CAST(ntt."id" AS VARCHAR(64)) AS "id", + CAST(ntt."id_profile" AS INTEGER) AS "id_profile", + TO_TIMESTAMP(ntt."date", 'DD.MM.YYYY"T"HH24:MI:SS') AS "date", + CAST(ntt."amount" AS INTEGER) AS "amount" + FROM "non_typed_table" AS ntt; +``` + +**Step 3: Run the Transformation** +Execute the transformation and wait for it to complete. + +**Step 4: Verify the Schema** +Once the transformation is finished, check the schema of the newly created table, `typed_table`. It should now include the appropriate data types. + +***Note:** [Incremental loading](/storage/tables/incremental-loading/) cannot be used when creating a typed table in this manner.* + +## Incremental Loading +The behavior of incremental loading differs between **typed** and **non-typed tables**: + +- **Typed tables:** Only the columns in the table's **primary key** are compared to detect changes. +- **Non-typed tables:** The entire row is compared, and rows are updated if **any value** has changed. + +For more information, refer to our documentation on [incremental loading](/storage/tables/incremental-loading/#difference-between-tables-with-native-datatypes-and-string-tables). + +## Handling NULLs +Data can contain `NULL` values or empty strings, which are converted differently based on the processing backend. + +### CSV Format: Unquoted vs Quoted Empty Values +In CSV files, there are two ways to represent an empty value: + +- **Unquoted empty** (`,,`) — the field is completely absent between delimiters +- **Quoted empty** (`""`) — the field explicitly contains an empty string + +This distinction matters because backends treat them differently. + +### Backend Behavior + +| CSV value | Snowflake (STRING) | BigQuery (STRING) | BigQuery (INT64/NUMERIC) | +|---|---|---|---| +| `,,` (unquoted empty) | `NULL` | `NULL` | `NULL` | +| `""` (quoted empty) | `NULL` | `’’` (empty string) | `NULL` | + +**Key difference:** Snowflake converts both `,,` and `""` to `NULL` for string columns. BigQuery preserves `""` as an empty string for STRING columns but coerces it to `NULL` for numeric columns (since an empty string is not a valid number). + +This means the **same CSV file** can produce different NULL semantics depending on the target column type in BigQuery. Numeric columns will have `NULL` where string columns will have an empty string, even though the source CSV value is the same. + +**Note:** This is a behavior of BigQuery's CSV parser, not Keboola Storage. BigQuery does not provide any option (similar to Snowflake's `EMPTY_FIELD_AS_NULL`) to change how quoted empty strings are handled during CSV loading. There is no available workaround at the BigQuery level. + +### The `treatValuesAsNull` Parameter +When importing data, you can use the `treatValuesAsNull` parameter to specify which values should be treated as `NULL`. For example, `treatValuesAsNull: [""]` tells the system to treat empty strings as `NULL`. + +However, in BigQuery, this parameter maps to BigQuery’s native [`nullMarker`](https://cloud.google.com/bigquery/docs/loading-data-cloud-storage-csv#csv-options) CSV load option, which **only applies to unquoted CSV fields**. Quoted values bypass the null marker check entirely — this is a limitation of BigQuery’s CSV parser, not Keboola Storage. + +In practice: + +| CSV value | `treatValuesAsNull: [""]` | BigQuery STRING result | +|---|---|---| +| `,,` (unquoted empty) | matches null marker → **NULL** | `NULL` | +| `""` (quoted empty) | skipped (quoted) | `’’` (empty string) | + +If your data source (e.g., a database extractor) writes `NULL` values as `""` instead of `,,`, string columns will contain empty strings instead of `NULL` in BigQuery, while numeric columns will still show `NULL` due to type coercion. To ensure correct NULL handling, the data source must represent NULL values as unquoted empty fields (`,,`) in the CSV. + +### NULL and Incremental Loading +Columns without native types are always `VARCHAR NOT NULL`. This means you don’t need to worry about specific `NULL` behavior. However, this changes with typed columns. + +In most databases, `NULL` does not equal `NULL` (`NULL == NULL` is not `TRUE`, but `NULL`). This behavior can disrupt the incremental loading process, where columns are compared to detect changes. + +To avoid such issues, ensure that your primary key columns are **not nullable**. This is especially relevant in `CTAS` (Create Table As Select) queries, where columns are nullable by default. To address this, explicitly define the columns as non-nullable in the `CTAS` expression. For example: + +``` +CREATE TABLE "ctas_table" ( + "id" NUMBER NOT NULL, + "name" VARCHAR(255) NOT NULL, + "created_at" TIMESTAMP_NTZ NOT NULL +) AS SELECT * FROM "typed_table"; +``` + diff --git a/src/content/docs/storage/tables/incremental-loading/index.md b/src/content/docs/storage/tables/incremental-loading/index.md new file mode 100644 index 000000000..ddf25c9fa --- /dev/null +++ b/src/content/docs/storage/tables/incremental-loading/index.md @@ -0,0 +1,311 @@ +--- +title: Primary Keys & Incremental Loading +slug: 'storage/tables/incremental-loading' +description: How Keboola loads data into tables — primary keys and deduplication, incremental loads, the _timestamp column with native data types, and automatic and manual incremental processing. +--- + +How data lands in a [table](/storage/tables/) depends on two settings: whether the table has a **primary key**, and whether the load is **incremental**. This page explains both, plus **incremental processing** — letting components read only the rows that changed since their last run. + +## Primary Keys +Each table may have a **primary key** defined on one or more columns. A primary key represents an +identifier of each row in the table. Each primary key can be defined manually on a table or as part of +[output mapping](/transformations/mappings/#output-mapping) of +[transformations](/transformations/) and [applications](/components/applications/). +The settings on both places must match, otherwise you will receive an error: + + Output mapping does not match destination table: primary key '' does not match 'Id' in 'out.c-tutorial.opportunity_denorm' (check transformations Denormalize opportunities (id opportunity.denormalize-opportunities)). + +This means that you cannot change the primary key of a table arbitrarily. Also note that you cannot set +the primary key on a column which contains duplicates --- you will receive the following error: + + Cannot create new primary key, duplicate values in primary key columns + +If you want to manually set a primary key on a table, you can do so in **Storage**: + +![Screenshot - Create Primary Key](/storage/tables/create-primary-key-1.png) + +Then select the columns you wish to add to the primary key: + +![Screenshot - Select columns](/storage/tables/create-primary-key-2.png) + +To remove an existing primary key, click the **bin** icon: + +![Screenshot - Remove Primary Key](/storage/tables/remove-primary-key.png) + +***Note:** Be aware that creating or removing the primary key can take some time on large tables.* + +### Primary Key Deduplication +When a primary key is defined on a column, the value of that column is guaranteed to be **unique** in that table. +As data is loaded into the table, only one of the rows with duplicate values is preserved. +All the other duplicates are ignored. +Let's say you have a table with two columns: `name` and `money`. The primary key is defined +on the column `name`. + +|name|money| +|---|---| +|John|$150| +|John|$340| +|Darla|$600| +|Annie|$500| +|John|$340000| +|Darla|$600000| + +Their uniqueness is checked and the data are de-duplicated. The result table looks like this: + +|name|money| +|---|---| +|Darla|$600| +|John|$340000| +|Annie|$500| + +The order of rows in the imported file is not important and is not kept. That means that from each of +the duplicate rows a randomly selected one is kept and all others are discarded. +In our example, the rows `John,$150`, `John,$340` and `Darla,$600000` were discarded. + +With a primary key defined on **multiple columns**, the combination of their values is unique. +Let's say you have a table with three columns: `name`, `age` and `money`. The primary key is defined +on two of them: `name` and `age`. +When you load the following data into your table: + +|name|age|money| +|---|---|---| +|John|15|$150| +|John|34|$340| +|Darla|60|$600| +|Annie|30|$500| +|John|34|$340000| +|Darla|60|$600000| + +their uniqueness is checked and the data are de-duplicated. The result table looks like this: + +|name|age|money| +|---|---|---| +|John|15|$150| +|Darla|60|$600| +|John|34|$340000| +|Annie|30|$500| + +Again, the order of rows in the imported file is not important and is not kept. +In our example, the rows `John,34,$340` and `Darla,60,$600000` were discarded. + +### Incremental Loading +When a primary key is defined on a column, it is also possible to take advantage of incremental loads. +If you load data into a table incrementally, new rows will be added and existing rows will be updated +unless they are completely identical to the existing rows. No rows will be deleted. +If you have a table with a primary key defined on the column `name`: + +|name|money| +|---|---| +|John|$150| +|Peter|$340| +|Darla|$600| + +and you import the following data to the table: + +|name|money| +|---|---| +|Annie|$500000| +|Peter|$340000| +|Darla|$600000| + +the result table will contain: + +|name|money| +|---|---| +|John|$150| +|Darla|$600000| +|Peter|$340000| +|Annie|$500000| + +When importing data into a table with a primary key, the uniqueness is checked. +The record `Peter,$340000` will overwrite the row `Peter,$340`, because it has the same primary key value. +The above applies only when **incremental load** is used. + +When an incremental load is not used, the contents of the target table are cleared before the load. When a primary key +is not defined and an incremental load is used, it simply appends the data to the table and does not update anything. + +#### Difference between tables with [native datatypes](/storage/tables/data-types/) and string tables + + +There is significant change when loading incrementally into table with native datatypes on. If a table does not have native datatypes enabled during incremental loading, the `_timestamp` column is updated based on the primary key only when a value in the row changes. In tables with native datatypes, the `_timestamp` column is updated every time when duplicate primary keys are imported. This behavior has an impact on [incremental processing](#incremental-processing). When rows with duplicate primary keys are imported into tables with native types, they are treated as new rows. + +**Example:** + +- Keboola Storage table newly created at **Tue Nov 22 2022 15:37:19 GMT+0000 (1669131439)** + +|ID|NAME|SKU|VALUE|DATE|_timestamp| +|---|---|---|---|---|---| +|1|John|CD-CZ-01|9247|2005-12-11|1669131439| +|2|Jack|CE-CA-22|3544|2012-10-14|1669131439| +|3|Jim|ED-BT-13|5262|2001-04-20|1669131439| +|4|Jil|BA-AB-11|5278|2014-12-14|1669131439| + +- Incremental import A1 at **Wed Nov 23 2022 16:41:20 GMT+0000 (1669221680)** + +| | ID | NAME | SKU | VALUE | DATE | +|------------|----|------|----------|-------|------------| +| new row => | 5 | Andy | AB-CF-48 | 7081 | 2003-07-05 | +| new row => | 6 | Beth | HH-FR-14 | 7541 | 2002-04-01 | + +- Result of incremental import A1 + +| |ID|NAME|SKU|VALUE|DATE| _timestamp | +|-----------------------------------|---|---|---|---|---|----------------| +| |1|John|CD-CZ-01|9247|2005-12-11| 1669131439 | +| |2|Jack|CE-CA-22|3544|2012-10-14| 1669131439 | +| |3|Jim|ED-BT-13|5262|2001-04-20| 1669131439 | +| |4|Jil|BA-AB-11|5278|2014-12-14| 1669131439 | +| added row = new _timestamp => |5|Andy|AB-CF-48|7081|2003-07-05| **1669221680** | +| added row = new _timestamp => |6|Beth|HH-FR-14|7541|2002-04-01| **1669221680** | + +- Incremental import A2 at **Wed Nov 23 2022 16:42:42 GMT+0000 (1669221762)** + +| | ID |NAME|SKU|VALUE|DATE| +|----------------------------------|---|---|---|---|---| +| existing row, no new values => |5|Andy|AB-CF-48|7081|2003-07-05| +| new row => |7|Edith|ED-BT-13|9471|1996-12-18| + +- Result of incremental import A2 + +Here we can see a **significant change in the incremental load**. The `_timestamp` column is updated for row `id:5`. For tables without native types, the row would not have the new value of `_timestamp`. + +| |ID|NAME|SKU|VALUE|DATE| _timestamp | +|--------------------------------------|---|---|---|---|---|----------------| +| | 1 |John|CD-CZ-01|9247|2005-12-11| 1669131439 | +| | 2 |Jack|CE-CA-22|3544|2012-10-14| 1669131439 | +| | 3 |Jim|ED-BT-13|5262|2001-04-20| 1669131439 | +| | 4 |Jil|BA-AB-11|5278|2014-12-14| 1669131439 | +| **updating row = new _timestamp =>** | 5 |Andy|AB-CF-48|7081|2003-07-05| **1669221762** | +| | 6 |Beth|HH-FR-14|7541|2002-04-01| 1669221680 | +| added row = new _timestamp => | 7 |Edith|ED-BT-13|9471|1996-12-18| **1669221762** | + +- Import A3 at **Wed Nov 23 2022 16:44:34 GMT+0000 (1669221874)** + +| | ID |NAME|SKU| VALUE |DATE| +|---------------------------------|---|---|---|----------------|---| +| existing row, with new value => |5|Andy|AB-CF-48| **6081** |2003-07-05| +| existing row, no new values => |7|Edith|ED-BT-13| 9471 |1996-12-18| +| new row => |8|Kate|CD-CZ-01| 5282 |2008-06-07| +| new row => |9|Josh|BA-AB-11| 6624 |2004-10-04| +| new row => |10|Arthur|EE-FF-66| 596 |2021-04-06 | + +- Result of incremental import A3 + +- Here we can see **another change that occurs only for tables with native types**. The `_timestamp` column for row `id:7` is updated but there was no change in it. + +| | ID |NAME|SKU| VALUE |DATE| _timestamp | +|--------------------------------------|-|---|---|----------|---|----------------| +| |1|John|CD-CZ-01| 9247 |2005-12-11| 1669131439 | +| |2|Jack|CE-CA-22| 3544 |2012-10-14| 1669131439 | +| |3|Jim|ED-BT-13| 5262 |2001-04-20| 1669131439 | +| |4|Jil|BA-AB-11| 5278 |2014-12-14| 1669131439 | +| updating row = new _timestamp => |5|Andy|AB-CF-48| **6081** |2003-07-05| **1669221874** | +| |6| Beth |HH-FR-14| 7541 |2002-04-01| 1669221680 | +| **updating row = new _timestamp =>** |7|Edith|ED-BT-13| 9471 |1996-12-18| **1669221874** | +| added row = new _timestamp => |8|Kate|CD-CZ-01| 5282 |2008-06-07| **1669221874** | +| added row = new _timestamp => |9|Josh|BA-AB-11| 6624 |2004-10-04| **1669221874** | +| added row = new _timestamp => |10|Arthur|EE-FF-66| 596 |2021-04-06| **1669221874** | + +### Incremental Processing +When a table is loaded incrementally, the update time of each row is recorded internally. This information +can be later used in input mapping of many components (especially [transformations](/transformations/mappings/#input-mapping)). +Incremental processing is available in two flavors --- *automatic* and *manual*. Incremental processing makes sense only for +components reading data from the Storage (e.g., transformations and writers). Note that this is not supported for all components yet. + +#### Automatic incremental processing +With automatic incremental processing, the component will receive only data modified from the last **successful** run of that component. +Extending the [above example](#incremental-loading) --- if you have a table with a primary key defined on the column `name`: + +|name|money| +|---|---| +|John|$150| +|Peter|$340| +|Darla|$600| + +and you import the following data to the table: + +|name|money| +|---|---| +|Darla|$600000| +|Peter|$340| +|Annie|$500000| +|Melanie|$900000| + +the result table will contain: + +|name|money| +|---|---| +|John|$150| +|**Darla**|**$600000**| +|Peter|$340| +|**Annie**|**$500000**| +|**Melanie**|**$900000**| + +Notice that the record for *Peter* was **not updated** because it was not changed at all (the +imported row was completely identical to the existing one). Therefore, +when using incremental processing, that row will not be loaded in input mapping. + +The component (e.g., writer) will receive only the highlighted rows. If there are no added or updated rows since the last successful run, +an empty table will be passed to the component. The image below shows the setting of automatic incremental processing for +the [Snowflake writer](/components/writers/database/snowflake/): + +![Screenshot - Automatic Incremental Processing](/storage/tables/adaptive-1.png) + +![Screenshot - Automatic Incremental Processing Detail](/storage/tables/adaptive-2.png) + +Automatic incremental processing offers the most efficient way to process data, but it is less transparent compared to processing the full tables. +The date used to identify newly arrived data is stored internally and can be reset via the **Reset** button. This will cause the +component to process the entire input table on the next run. + +You can always verify the load date used in a particular job using the [Job events](/management/jobs/). Search for an event *Exported table X*: + +![Screenshot - Automatic Incremental Processing Events](/storage/tables/adaptive-events-1.png) + +Click the event to show its detail; the `changedSince` parameter shows the date used to select the added and updated data: + +![Screenshot - Automatic Incremental Processing Event Detail](/storage/tables/adaptive-events-2.png) + +#### Manual incremental processing +When the [Automatic Incremental Processing](#automatic-incremental-processing) feature is not available, or you +need greater control of the processing, you may specify the `Changed in last` option. +This enables the component to process only a specified increment of the data: + +![Screenshot - Incremental Processing](/storage/tables/incremental-processing.png) + +Extending the [above example](#incremental-loading) --- if you have a table with a primary key defined on the column `name`: + +|name|money| +|---|---| +|John|$150| +|Peter|$340| +|Darla|$600| + +and you import the following data to the table: + +|name|money| +|---|---| +|Darla|$600000| +|Peter|$340| +|Annie|$500000| +|Melanie|$900000| + +assuming that the import was on 2010-01-02 10:00, the result table will contain (the *\*_timestamp\** column is not an actual column of the table, it is just displayed here for illustration purposes): + +|name|money|*\*_timestamp\** +|---|---|---| +|John|$150|2010-01-01 10:00| +|Darla|$600000|**2010-01-02 10:00**| +|Peter|$340|2010-01-01 10:00| +|Annie|$500000|**2010-01-02 10:00**| +|Melanie|$900000|**2010-01-02 10:00**| + +Therefore, three rows in the table will be considered as changed (either added or updated). Now when you run a component (e.g., transformation) +with the `Changed in last` option, various things can happen: + +- `Changed in last` is set to `1 day`, and the component is started any time between `2010-01-02 10:00` and `2010-01-03 10:00` -- **three** rows will be exported from the table. +- `Changed in last` is set to `1 day`, and the component is started any time after `2010-01-03 10:00` -- **no** rows will be exported from the table. +- `Changed in last` is set to `2 day`, and the component is started any time between `2010-01-02 10:00` and `2010-01-03 10:00` -- **five** rows will be exported from the table. + +Notice that the record for *Peter* was **not updated** because it did not change (the +imported row was completely identical to the existing one). Therefore, +when using incremental processing set to `1 day`, that row will not be included in the input mapping. diff --git a/src/content/docs/storage/tables/index.md b/src/content/docs/storage/tables/index.md index 567a0c346..d0165f9b1 100644 --- a/src/content/docs/storage/tables/index.md +++ b/src/content/docs/storage/tables/index.md @@ -1,413 +1,117 @@ ---- -title: Tables -slug: 'storage/tables' ---- - - - -The *Table Storage* for your project is available under the **Tables** tab in the Storage section. -All data tables are organized into [buckets](/storage/buckets/) that can also be -used to [share tables](/catalog/) between projects. - -The actual data tables and the buckets are created primarily by Keboola components (data source connectors, transformations, -and applications), or they are imported from CSV files. If you want to import data into an already -**existing table**, the imported table must contain all the columns of the existing table, even if the existing -table is empty. If any columns are missing, you will receive an error message similar to the following: - - Some columns are missing in the CSV file. Missing columns: lat,long. Expected columns: lat,long. - Please check if the expected "," delimiter is used in the CSV file. - -The imported file **may** also contain additional columns not present in the existing -table. In that case, the columns from the imported table will be added to the existing table. - -Table and column names can contain only alphanumeric characters. Dash and -underscores are allowed. Column names must not start or end with dash `-` or underscore -character `_`. - -When you select a table from any bucket in Storage, detailed information about the table will be displayed -at the top of the screen. This is what we refer to as the **table detail** throughout our documentation. - -## Aliases -Apart from actual tables, it is also possible to create aliases. They behave similarly to -[database views](https://en.wikipedia.org/wiki/View_(SQL)). - -An alias does not contain any actual data; it is simply a link to some already existing data. -Therefore, an alias cannot be written to, and its size does not count toward your project quota. - -Alias tables are automatically materialized as physical database VIEWs. This makes them fully accessible in workspaces and -transformations via [read-only storage access](/transformations/mappings/#read-only-input-mapping) — no input mapping configuration is required. -Filtered aliases are also supported; the filter condition is enforced as a `WHERE` clause in the VIEW. -In [linked buckets](/catalog/), alias VIEWs from the source project are automatically mirrored to the destination project, -making them immediately queryable there as well. - -To create an alias table, go to the table detail, click the three dots on the right side of the screen, and select the 'Create alias table' option. - -![Screenshot - Create alias](/storage/tables/create-alias.png) - -An alias table can be filtered by a simple condition. - -![Screenshot - Create Simple alias](/storage/tables/create-simple-alias.png) - -The table detail of an alias table contains additional information, including a reference to the source table from which it was created and any filters applied to the alias. Note that you can adjust the alias filters even after the alias table has been created. - -![Screenshot - Simple alias result](/storage/tables/create-simple-alias-result.png) - -When attempting to delete a table with alias tables elsewhere in Storage, you will be prompted with a notification as part of the deletion process. The notification will detail the aliases (including links) connected to the table. You must confirm that you understand the aliases will be deleted as well before proceeding. - -![Screenshot - Deleting table having aliases](/storage/tables/delete-table-with-alias.png) - -By default, alias columns are automatically synchronized with the source table. Columns added to the source -table will be added to the alias automatically. -You can prevent this by disabling *Synchronize columns with source table*. - -Aliases with automatically synchronized columns and without a filter can be chained. - -## Metadata -Each [Table Storage](/storage/) object (bucket, table, column) has an associated key-value -store. This can be used to store arbitrary metadata (information about the data itself). Apart from -arbitrary user-defined metadata, there is also some information stored automatically. For example, -each bucket and table has information about which configuration of which component created them. -One important use case of metadata is [Column Data Types](/storage/tables/data-types/). - -## Primary Keys -Each table may have a **primary key** defined on one or more columns. A primary key represents an -identifier of each row in the table. Each primary key can be defined manually on a table or as part of -[output mapping](/transformations/mappings/#output-mapping) of -[transformations](/transformations/) and [applications](/components/applications/). -The settings on both places must match, otherwise you will receive an error: - - Output mapping does not match destination table: primary key '' does not match 'Id' in 'out.c-tutorial.opportunity_denorm' (check transformations Denormalize opportunities (id opportunity.denormalize-opportunities)). - -This means that you cannot change the primary key of a table arbitrarily. Also note that you cannot set -the primary key on a column which contains duplicates — you will receive the following error: - - Cannot create new primary key, duplicate values in primary key columns - -If you want to manually set a primary key on a table, you can do so in **Storage**: - -![Screenshot - Create Primary Key](/storage/tables/create-primary-key-1.png) - -Then select the columns you wish to add to the primary key: - -![Screenshot - Select columns](/storage/tables/create-primary-key-2.png) - -To remove an existing primary key, click the **bin** icon: - -![Screenshot - Remove Primary Key](/storage/tables/remove-primary-key.png) - -***Note:** Be aware that creating or removing the primary key can take some time on large tables.* - -### Primary Key Deduplication -When a primary key is defined on a column, the value of that column is guaranteed to be **unique** in that table. -As data is loaded into the table, only one of the rows with duplicate values is preserved. -All the other duplicates are ignored. -Let's say you have a table with two columns: `name` and `money`. The primary key is defined -on the column `name`. - -|name|money| -|---|---| -|John|$150| -|John|$340| -|Darla|$600| -|Annie|$500| -|John|$340000| -|Darla|$600000| - -Their uniqueness is checked and the data are de-duplicated. The result table looks like this: - -|name|money| -|---|---| -|Darla|$600| -|John|$340000| -|Annie|$500| - -The order of rows in the imported file is not important and is not kept. That means that from each of -the duplicate rows a randomly selected one is kept and all others are discarded. -In our example, the rows `John,$150`, `John,$340` and `Darla,$600000` were discarded. - -With a primary key defined on **multiple columns**, the combination of their values is unique. -Let's say you have a table with three columns: `name`, `age` and `money`. The primary key is defined -on two of them: `name` and `age`. -When you load the following data into your table: - -|name|age|money| -|---|---|---| -|John|15|$150| -|John|34|$340| -|Darla|60|$600| -|Annie|30|$500| -|John|34|$340000| -|Darla|60|$600000| - -their uniqueness is checked and the data are de-duplicated. The result table looks like this: - -|name|age|money| -|---|---|---| -|John|15|$150| -|Darla|60|$600| -|John|34|$340000| -|Annie|30|$500| - -Again, the order of rows in the imported file is not important and is not kept. -In our example, the rows `John,34,$340` and `Darla,60,$600000` were discarded. - -### Incremental Loading -When a primary key is defined on a column, it is also possible to take advantage of incremental loads. -If you load data into a table incrementally, new rows will be added and existing rows will be updated -unless they are completely identical to the existing rows. No rows will be deleted. -If you have a table with a primary key defined on the column `name`: - -|name|money| -|---|---| -|John|$150| -|Peter|$340| -|Darla|$600| - -and you import the following data to the table: - -|name|money| -|---|---| -|Annie|$500000| -|Peter|$340000| -|Darla|$600000| - -the result table will contain: - -|name|money| -|---|---| -|John|$150| -|Darla|$600000| -|Peter|$340000| -|Annie|$500000| - -When importing data into a table with a primary key, the uniqueness is checked. -The record `Peter,$340000` will overwrite the row `Peter,$340`, because it has the same primary key value. -The above applies only when **incremental load** is used. - -When an incremental load is not used, the contents of the target table are cleared before the load. When a primary key -is not defined and an incremental load is used, it simply appends the data to the table and does not update anything. - -#### Difference between tables with [native datatypes](/storage/tables/data-types/) and string tables - -There is significant change when loading incrementally into table with native datatypes on. If a table does not have native datatypes enabled during incremental loading, the `_timestamp` column is updated based on the primary key only when a value in the row changes. In tables with native datatypes, the `_timestamp` column is updated every time when duplicate primary keys are imported. This behavior has an impact on [incremental processing](/storage/tables/#incremental-processing). When rows with duplicate primary keys are imported into tables with native types, they are treated as new rows. - -**Example:** - -- Keboola Storage table newly created at **Tue Nov 22 2022 15:37:19 GMT+0000 (1669131439)** - -|ID|NAME|SKU|VALUE|DATE|_timestamp| -|---|---|---|---|---|---| -|1|John|CD-CZ-01|9247|2005-12-11|1669131439| -|2|Jack|CE-CA-22|3544|2012-10-14|1669131439| -|3|Jim|ED-BT-13|5262|2001-04-20|1669131439| -|4|Jil|BA-AB-11|5278|2014-12-14|1669131439| - -- Incremental import A1 at **Wed Nov 23 2022 16:41:20 GMT+0000 (1669221680)** - -| | ID | NAME | SKU | VALUE | DATE | -|------------|----|------|----------|-------|------------| -| new row => | 5 | Andy | AB-CF-48 | 7081 | 2003-07-05 | -| new row => | 6 | Beth | HH-FR-14 | 7541 | 2002-04-01 | - -- Result of incremental import A1 - -| |ID|NAME|SKU|VALUE|DATE| _timestamp | -|-----------------------------------|---|---|---|---|---|----------------| -| |1|John|CD-CZ-01|9247|2005-12-11| 1669131439 | -| |2|Jack|CE-CA-22|3544|2012-10-14| 1669131439 | -| |3|Jim|ED-BT-13|5262|2001-04-20| 1669131439 | -| |4|Jil|BA-AB-11|5278|2014-12-14| 1669131439 | -| added row = new _timestamp => |5|Andy|AB-CF-48|7081|2003-07-05| **1669221680** | -| added row = new _timestamp => |6|Beth|HH-FR-14|7541|2002-04-01| **1669221680** | - -- Incremental import A2 at **Wed Nov 23 2022 16:42:42 GMT+0000 (1669221762)** - -| | ID |NAME|SKU|VALUE|DATE| -|----------------------------------|---|---|---|---|---| -| existing row, no new values => |5|Andy|AB-CF-48|7081|2003-07-05| -| new row => |7|Edith|ED-BT-13|9471|1996-12-18| - -- Result of incremental import A2 - -Here we can see a **significant change in the incremental load**. The `_timestamp` column is updated for row `id:5`. For tables without native types, the row would not have the new value of `_timestamp`. - -| |ID|NAME|SKU|VALUE|DATE| _timestamp | -|--------------------------------------|---|---|---|---|---|----------------| -| | 1 |John|CD-CZ-01|9247|2005-12-11| 1669131439 | -| | 2 |Jack|CE-CA-22|3544|2012-10-14| 1669131439 | -| | 3 |Jim|ED-BT-13|5262|2001-04-20| 1669131439 | -| | 4 |Jil|BA-AB-11|5278|2014-12-14| 1669131439 | -| **updating row = new _timestamp =>** | 5 |Andy|AB-CF-48|7081|2003-07-05| **1669221762** | -| | 6 |Beth|HH-FR-14|7541|2002-04-01| 1669221680 | -| added row = new _timestamp => | 7 |Edith|ED-BT-13|9471|1996-12-18| **1669221762** | - -- Import A3 at **Wed Nov 23 2022 16:44:34 GMT+0000 (1669221874)** - -| | ID |NAME|SKU| VALUE |DATE| -|---------------------------------|---|---|---|----------------|---| -| existing row, with new value => |5|Andy|AB-CF-48| **6081** |2003-07-05| -| existing row, no new values => |7|Edith|ED-BT-13| 9471 |1996-12-18| -| new row => |8|Kate|CD-CZ-01| 5282 |2008-06-07| -| new row => |9|Josh|BA-AB-11| 6624 |2004-10-04| -| new row => |10|Arthur|EE-FF-66| 596 |2021-04-06 | - -- Result of incremental import A3 - -- Here we can see **another change that occurs only for tables with native types**. The `_timestamp` column for row `id:7` is updated but there was no change in it. - -| | ID |NAME|SKU| VALUE |DATE| _timestamp | -|--------------------------------------|-|---|---|----------|---|----------------| -| |1|John|CD-CZ-01| 9247 |2005-12-11| 1669131439 | -| |2|Jack|CE-CA-22| 3544 |2012-10-14| 1669131439 | -| |3|Jim|ED-BT-13| 5262 |2001-04-20| 1669131439 | -| |4|Jil|BA-AB-11| 5278 |2014-12-14| 1669131439 | -| updating row = new _timestamp => |5|Andy|AB-CF-48| **6081** |2003-07-05| **1669221874** | -| |6| Beth |HH-FR-14| 7541 |2002-04-01| 1669221680 | -| **updating row = new _timestamp =>** |7|Edith|ED-BT-13| 9471 |1996-12-18| **1669221874** | -| added row = new _timestamp => |8|Kate|CD-CZ-01| 5282 |2008-06-07| **1669221874** | -| added row = new _timestamp => |9|Josh|BA-AB-11| 6624 |2004-10-04| **1669221874** | -| added row = new _timestamp => |10|Arthur|EE-FF-66| 596 |2021-04-06| **1669221874** | - -### Incremental Processing -When a table is loaded incrementally, the update time of each row is recorded internally. This information -can be later used in input mapping of many components (especially [transformations](/transformations/mappings/#input-mapping)). -Incremental processing is available in two flavors — *automatic* and *manual*. Incremental processing makes sense only for -components reading data from the Storage (e.g., transformations and writers). Note that this is not supported for all components yet. - -#### Automatic incremental processing -With automatic incremental processing, the component will receive only data modified from the last **successful** run of that component. -Extending the [above example](#incremental-loading) — if you have a table with a primary key defined on the column `name`: - -|name|money| -|---|---| -|John|$150| -|Peter|$340| -|Darla|$600| - -and you import the following data to the table: - -|name|money| -|---|---| -|Darla|$600000| -|Peter|$340| -|Annie|$500000| -|Melanie|$900000| - -the result table will contain: - -|name|money| -|---|---| -|John|$150| -|**Darla**|**$600000**| -|Peter|$340| -|**Annie**|**$500000**| -|**Melanie**|**$900000**| - -Notice that the record for *Peter* was **not updated** because it was not changed at all (the -imported row was completely identical to the existing one). Therefore, -when using incremental processing, that row will not be loaded in input mapping. - -The component (e.g., writer) will receive only the highlighted rows. If there are no added or updated rows since the last successful run, -an empty table will be passed to the component. The image below shows the setting of automatic incremental processing for -the [Snowflake writer](/components/writers/database/snowflake/): - -![Screenshot - Automatic Incremental Processing](/storage/tables/adaptive-1.png) - -![Screenshot - Automatic Incremental Processing Detail](/storage/tables/adaptive-2.png) - -Automatic incremental processing offers the most efficient way to process data, but it is less transparent compared to processing the full tables. -The date used to identify newly arrived data is stored internally and can be reset via the **Reset** button. This will cause the -component to process the entire input table on the next run. - -You can always verify the load date used in a particular job using the [Job events](/management/jobs/). Search for an event *Exported table X*: - -![Screenshot - Automatic Incremental Processing Events](/storage/tables/adaptive-events-1.png) - -Click the event to show its detail; the `changedSince` parameter shows the date used to select the added and updated data: - -![Screenshot - Automatic Incremental Processing Event Detail](/storage/tables/adaptive-events-2.png) - -#### Manual incremental processing -When the [Automatic Incremental Processing](#automatic-incremental-processing) feature is not available, or you -need greater control of the processing, you may specify the `Changed in last` option. -This enables the component to process only a specified increment of the data: - -![Screenshot - Incremental Processing](/storage/tables/incremental-processing.png) - -Extending the [above example](#incremental-loading) — if you have a table with a primary key defined on the column `name`: - -|name|money| -|---|---| -|John|$150| -|Peter|$340| -|Darla|$600| - -and you import the following data to the table: - -|name|money| -|---|---| -|Darla|$600000| -|Peter|$340| -|Annie|$500000| -|Melanie|$900000| - -assuming that the import was on 2010-01-02 10:00, the result table will contain (the *\*_timestamp\** column is not an actual column of the table, it is just displayed here for illustration purposes): - -|name|money|*\*_timestamp\** -|---|---|---| -|John|$150|2010-01-01 10:00| -|Darla|$600000|**2010-01-02 10:00**| -|Peter|$340|2010-01-01 10:00| -|Annie|$500000|**2010-01-02 10:00**| -|Melanie|$900000|**2010-01-02 10:00**| - -Therefore, three rows in the table will be considered as changed (either added or updated). Now when you run a component (e.g., transformation) -with the `Changed in last` option, various things can happen: - -- `Changed in last` is set to `1 day`, and the component is started any time between `2010-01-02 10:00` and `2010-01-03 10:00` -- **three** rows will be exported from the table. -- `Changed in last` is set to `1 day`, and the component is started any time after `2010-01-03 10:00` -- **no** rows will be exported from the table. -- `Changed in last` is set to `2 day`, and the component is started any time between `2010-01-02 10:00` and `2010-01-03 10:00` -- **five** rows will be exported from the table. - -Notice that the record for *Peter* was **not updated** because it did not change (the -imported row was completely identical to the existing one). Therefore, -when using incremental processing set to `1 day`, that row will not be included in the input mapping. - -## Truncate Table -**Overview** - -The Truncate Table feature removes all records from a table while preserving its schema (column structure and metadata). The feature deletes all rows but keeps the table structure intact. Primary keys or metadata remain without change. The feature is highly efficient especially when replacing an entire dataset. - -**When to Use Truncate Table?** -- Replacing data completely (e.g., daily refresh of customer records). -- Ensuring clean data ingestion without duplication. -- Automating scheduled table resets (e.g., overwriting temporary datasets). - -**How to Truncate a Table:** -1. Navigate to Storage → Tables. -2. Select the table you want to truncate. -3. Click the three dots on the right side of the screen. A drop-down menu will display. -4. Select "Truncate table". A warning information will display. -5. Confirm clicking the "Truncate table" button. - -## Delete Rows -**Overview** - -The Delete Rows feature allows you to remove specific records from a table in Storage. Unlike Truncate Table, which clears all data, Delete Rows selectively removes only matching records, keeping the rest of the table intact. - -When using the Delete Rows feature, you can apply multiple filters in a single execution, such as: -- Rows changed within a specific period. -- Rows containing a specific value in a given column. -Before deleting the rows, Keboola displays a preview of the records that will be removed based on your filters. - -**When to Use Delete Rows?** -- Removing outdated records (e.g., inactive customers). -- Correcting erroneous data entries. -- Synchronizing data with external systems (e.g., deleting stale records). - -**How to Delete Rows:** -1. Navigate to Storage → Tables. -2. Select the table where you want to delete rows. -3. Click the three dots on the right side of the screen to open a drop-down menu. -4. Select "Delete Rows". A pop-up will appear. -5. Set up the filters to specify which rows should be deleted. -6. Click "Delete rows" to confirm. +--- +title: Tables +slug: 'storage/tables' +description: Working with tables in Keboola Storage — importing data, table and column naming rules, alias tables, metadata, and the Truncate Table and Delete Rows features. +--- + + + +The *Table Storage* for your project is available under the **Tables & Buckets** tab in the Storage section. +All data tables are organized into [buckets](/storage/buckets/) that can also be +used to [share tables](/catalog/) between projects. + +The actual data tables and the buckets are created primarily by Keboola components (data source connectors, transformations, +and applications), or they are imported from CSV files. If you want to import data into an already +**existing table**, the imported table must contain all the columns of the existing table, even if the existing +table is empty. If any columns are missing, you will receive an error message similar to the following: + + Some columns are missing in the CSV file. Missing columns: lat,long. Expected columns: lat,long. + Please check if the expected "," delimiter is used in the CSV file. + +The imported file **may** also contain additional columns not present in the existing +table. In that case, the columns from the imported table will be added to the existing table. + +Table and column names can contain only alphanumeric characters. Dash and +underscores are allowed. Column names must not start or end with dash `-` or underscore +character `_`. + +When you select a table from any bucket in Storage, detailed information about the table will be displayed +at the top of the screen. This is what we refer to as the **table detail** throughout our documentation. + +## Aliases +Apart from actual tables, it is also possible to create aliases. They behave similarly to +[database views](https://en.wikipedia.org/wiki/View_(SQL)). + +An alias does not contain any actual data; it is simply a link to some already existing data. +Therefore, an alias cannot be written to, and its size does not count toward your project quota. + +Alias tables are automatically materialized as physical database VIEWs. This makes them fully accessible in workspaces and +transformations via [read-only storage access](/transformations/mappings/#read-only-input-mapping) — no input mapping configuration is required. +Filtered aliases are also supported; the filter condition is enforced as a `WHERE` clause in the VIEW. +In [linked buckets](/catalog/), alias VIEWs from the source project are automatically mirrored to the destination project, +making them immediately queryable there as well. + +To create an alias table, go to the table detail, click the three dots on the right side of the screen, and select the 'Create alias table' option. + +![Screenshot - Create alias](/storage/tables/create-alias.png) + +An alias table can be filtered by a simple condition. + +![Screenshot - Create Simple alias](/storage/tables/create-simple-alias.png) + +The table detail of an alias table contains additional information, including a reference to the source table from which it was created and any filters applied to the alias. Note that you can adjust the alias filters even after the alias table has been created. + +![Screenshot - Simple alias result](/storage/tables/create-simple-alias-result.png) + +When attempting to delete a table with alias tables elsewhere in Storage, you will be prompted with a notification as part of the deletion process. The notification will detail the aliases (including links) connected to the table. You must confirm that you understand the aliases will be deleted as well before proceeding. + +![Screenshot - Deleting table having aliases](/storage/tables/delete-table-with-alias.png) + +By default, alias columns are automatically synchronized with the source table. Columns added to the source +table will be added to the alias automatically. +You can prevent this by disabling *Synchronize columns with source table*. + +Aliases with automatically synchronized columns and without a filter can be chained. + +## Metadata +Each [Table Storage](/storage/) object (bucket, table, column) has an associated key-value +store. This can be used to store arbitrary metadata (information about the data itself). Apart from +arbitrary user-defined metadata, there is also some information stored automatically. For example, +each bucket and table has information about which configuration of which component created them. +One important use case of metadata is [Column Data Types](/storage/tables/data-types/). + +## Primary Keys & Incremental Loading + +Each table may have a **primary key** that identifies its rows: values are deduplicated on load, and with **incremental loading** new rows are added and existing rows updated instead of the table being cleared. Components can then use **incremental processing** to read only the rows changed since their last run. + +These load behaviors have their own page — see [Primary Keys & Incremental Loading](/storage/tables/incremental-loading/). + +## Truncate Table +**Overview** + +The Truncate Table feature removes all records from a table while preserving its schema (column structure and metadata). The feature deletes all rows but keeps the table structure intact. Primary keys or metadata remain without change. The feature is highly efficient especially when replacing an entire dataset. + +**When to Use Truncate Table?** +- Replacing data completely (e.g., daily refresh of customer records). +- Ensuring clean data ingestion without duplication. +- Automating scheduled table resets (e.g., overwriting temporary datasets). + +**How to Truncate a Table:** +1. Navigate to **Storage → Tables & Buckets**. +2. Select the table you want to truncate. +3. Click the three dots on the right side of the screen. A drop-down menu will display. +4. Select "Truncate table". A warning information will display. +5. Confirm clicking the "Truncate table" button. + +## Delete Rows +**Overview** + +The Delete Rows feature allows you to remove specific records from a table in Storage. Unlike Truncate Table, which clears all data, Delete Rows selectively removes only matching records, keeping the rest of the table intact. + +When using the Delete Rows feature, you can apply multiple filters in a single execution, such as: +- Rows changed within a specific period. +- Rows containing a specific value in a given column. +Before deleting the rows, Keboola displays a preview of the records that will be removed based on your filters. + +**When to Use Delete Rows?** +- Removing outdated records (e.g., inactive customers). +- Correcting erroneous data entries. +- Synchronizing data with external systems (e.g., deleting stale records). + +**How to Delete Rows:** +1. Navigate to **Storage → Tables & Buckets**. +2. Select the table where you want to delete rows. +3. Click the three dots on the right side of the screen to open a drop-down menu. +4. Select **Delete rows**. A pop-up will appear. +5. Set up the filters to specify which rows should be deleted. +6. Click "Delete rows" to confirm. diff --git a/src/content/docs/storage/tables/uploads.md b/src/content/docs/storage/tables/uploads.md index 62c032d27..2f5d53e60 100644 --- a/src/content/docs/storage/tables/uploads.md +++ b/src/content/docs/storage/tables/uploads.md @@ -1,24 +1,25 @@ ---- -title: Table Import & Export -slug: 'storage/tables/uploads' ---- - -All tables imported to and exported from Storage go through [Files](/storage/files/). - -When a table is **imported** into Storage by any means (manually, through a data source connector, or as a result of running an application), -the CSV file is first stored in *Files* and only then imported to an actual table. -This means that the Storage Files contain a history of data uploaded to the Storage Tables. -It is useful mainly in the two following cases: - -1. [Reverting table](/storage/tables/) content to a particular imported version -2. Analyzing how something got into a table (useful mainly for incremental loads) - -Every time a table is **exported** from Storage, the process is reversed: first, a file is -created in *Files* and then it is actually downloaded from there. This does not apply when exporting -Storage tables manually though. -Beware, however, that due to the nature of database exports, the exported table may be **sliced** and require -[substantial effort to reconstruct](https://developers.keboola.com/integrate/storage/api/import-export/#working-with-sliced-files). -To make sure your tables are exported as merged files, always use the **Export** feature in -the **Action** tab of the table detail: - -![Screenshot - Export table](/storage/tables/table-export.png) +--- +title: Table Import & Export +slug: 'storage/tables/uploads' +description: How table imports and exports flow through File Storage, why exported tables may be sliced, and how to export a merged file from the table detail. +--- + +All tables imported to and exported from Storage go through [Files](/storage/files/). + +When a table is **imported** into Storage by any means (manually, through a data source connector, or as a result of running an application), +the CSV file is first stored in *Files* and only then imported to an actual table. +This means that the Storage Files contain a history of data uploaded to the Storage Tables. +It is useful mainly in the two following cases: + +1. [Reverting table](/storage/tables/) content to a particular imported version +2. Analyzing how something got into a table (useful mainly for incremental loads) + +Every time a table is **exported** from Storage, the process is reversed: first, a file is +created in *Files* and then it is actually downloaded from there. This does not apply when exporting +Storage tables manually though. +Beware, however, that due to the nature of database exports, the exported table may be **sliced** and require +[substantial effort to reconstruct](https://developers.keboola.com/integrate/storage/api/import-export/#working-with-sliced-files). +To make sure your tables are exported as merged files, always use the **Export table** option in +the three-dots menu of the table detail: + +![Screenshot - Export table](/storage/tables/table-export.png) diff --git a/src/content/docs/transformations/mappings/index.md b/src/content/docs/transformations/mappings/index.md index 42c1a776b..445d1ce14 100644 --- a/src/content/docs/transformations/mappings/index.md +++ b/src/content/docs/transformations/mappings/index.md @@ -83,7 +83,7 @@ idea to add the `.csv` extension. For **Database Staging**, the table name is au use an arbitrary name. It also means that they may be case sensitive, for instance, if the destination is in Snowflake staging. - **Columns** — Select specific columns if you do not want to import them all; this saves processing time for larger tables. -- **Changed in last** — If you use [incremental processing](/storage/tables/#incremental-processing), +- **Changed in last** — If you use [incremental processing](/storage/tables/incremental-loading/#incremental-processing), this comes in handy; import only rows changed or created within the selected time period. The supported time dimensions are `minutes`, `hours`, and `days`. - **Data filter** — Filter source rows to the rows that match this single-column multiple-values filter. @@ -212,8 +212,8 @@ A cloned table is an exact copy of the source table, including the `_timestamp` This column is used internally by Keboola for comparison with the value of the *Changed in last* filter. The column type differs by backend: `TIMESTAMP_NTZ(9)` on Snowflake and `TIMESTAMP` on BigQuery. -The value contains the [last change of the row](/storage/tables/#manual-incremental-processing). -You can use this column to set up [incremental processing](/storage/tables/#incremental-processing), +The value contains the [last change of the row](/storage/tables/incremental-loading/#manual-incremental-processing). +You can use this column to set up [incremental processing](/storage/tables/incremental-loading/#incremental-processing), i.e., to replace the role of the **Changed in Last** filter in the input mapping (which you can't use with a clone mapping). **Important: The _timestamp column cannot be imported back to Storage.** @@ -303,8 +303,8 @@ the results of your transformation (i.e., contents of the Output Mapping *source - **Incremental** — Check this option to make sure that in case the *Destination* table already exists, it is not overwritten, but resulting data is appended to it. However, any existing row having the same primary key as a new row will be replaced. See the description of -[incremental loading](/storage/tables/#incremental-loading) for a detailed explanation and examples. -- **Primary key** — The [primary key](/storage/tables/#primary-keys) of the destination table; if the table already exists, +[incremental loading](/storage/tables/incremental-loading/) for a detailed explanation and examples. +- **Primary key** — The [primary key](/storage/tables/incremental-loading/#primary-keys) of the destination table; if the table already exists, the primary key must match. Feel free to use a multi-column primary key. - **Deduplication Strategy** — This allows to switch in Snowflake transformations from the default load Upsert to Insert. Upsert option uses deduplication based on primary keys and ensures data quality. By switching to Insert, the load performs faster but skips deduplication and type casting - meaning you are responsible for uniqueness and correct data types. As this option is for high data maturity users, you can ask our support to enable this for your project as Deduplication Strategy is under feature flag. - **Delete rows** — When Incremental loading is enabled, you can delete specific rows from the destination table before importing new data into the Storage. This gives you precise control over incremental updates. There are 2 options you can use: diff --git a/src/sidebar.mjs b/src/sidebar.mjs index b0047ea12..5368af508 100644 --- a/src/sidebar.mjs +++ b/src/sidebar.mjs @@ -514,6 +514,7 @@ export const sidebar = [ collapsed: true, items: [ { label: "Overview", slug: "storage/tables" }, + { slug: "storage/tables/incremental-loading" }, { slug: "storage/tables/data-types" }, { slug: "storage/tables/csv-files" }, { slug: "storage/tables/uploads" },