diff --git a/_data/navigation.yml b/_data/navigation.yml index 0b48c2c22..5a45880c5 100644 --- a/_data/navigation.yml +++ b/_data/navigation.yml @@ -598,6 +598,26 @@ items: - url: /storage/byobq/ title: Bring Your Own BigQuery + - url: /storage/api/ + title: Storage API + items: + - url: /storage/api/configurations/ + title: Configurations + - url: /storage/api/import-export/ + title: Import & Export + - url: /storage/api/importer/ + title: API Importer + - url: /storage/api/tde-exporter/ + title: TDE Exporter + - url: /storage/api/clients/python-client/ + title: Python client + - url: /storage/api/clients/r-client/ + title: R client + - url: /storage/api/clients/php-client/ + title: PHP client + - url: /storage/api/clients/docker-cli/ + title: Docker CLI client + - url: /transformations/ title: Transformations items: diff --git a/public/storage/api/async-import-handling.svg b/public/storage/api/async-import-handling.svg new file mode 100644 index 000000000..777e93372 --- /dev/null +++ b/public/storage/api/async-import-handling.svg @@ -0,0 +1 @@ + \ No newline at end of file diff --git a/public/storage/api/new-table.csv b/public/storage/api/new-table.csv new file mode 100644 index 000000000..8dbb6c464 --- /dev/null +++ b/public/storage/api/new-table.csv @@ -0,0 +1,5 @@ +"id","secondCol" +"1","a" +"2","b" +"3","c" +"4","d" \ No newline at end of file diff --git a/src/content/docs/storage/api/clients/docker-cli/index.md b/src/content/docs/storage/api/clients/docker-cli/index.md new file mode 100644 index 000000000..da0f0c21a --- /dev/null +++ b/src/content/docs/storage/api/clients/docker-cli/index.md @@ -0,0 +1,112 @@ +--- +title: Storage Docker CLI Client +slug: 'storage/api/clients/docker-cli' +redirect_from: + - /integrate/storage/docker-cli-client/ + - /integrate/storage/php-cli-client/ +--- + + +The Storage API Docker command line interface (CLI) client is a portable command line client which provides +a simple implementation of [Storage API](https://api.keboola.com/?service=storage). +It runs on any platform which has Docker installed. + +Currently, the client implements + +- functions for exporting and importing tables; +- functions for creating and deleting buckets; and additionally, +- the [project backup feature](/management/project-export/). + +The client source is available in our [Github repository](https://github.com/keboola/storage-api-cli). +The client docker image is available in the [Quay repository](https://quay.io/repository/keboola/storage-api-cli?tab=tags). + +## Running in Docker +To print available commands: + +```bash +docker run quay.io/keboola/storage-api-cli:latest +``` + +The `latest` image tag always refers to the latest tagged version. + +## Running Phar + +PHAR (PHP Archive) is now deprecated, but there are still some older versions available. See the [repository documentation](https://github.com/keboola/storage-api-cli#running-phar-deprecated). + +### Example --- Creating a Table +To create a new table in Storage, use the `create-table` command. Provide the name of an +existing bucket, the name of the new table and a CSV file with the table's contents. + +To create the`new-table` table in the `in.c-main` bucket, use + +```bash +docker run --volume=$("pwd"):/data quay.io/keboola/storage-api-cli:latest create-table in.c-main new-table /data/new-table.csv --token=storage_token +``` + +or on Windows: + + docker run --volume=C:\Users\name\some-dir:/data quay.io/keboola/storage-api-cli:latest create-table in.c-main new-table /data/new-table.csv --token=storage_token + +or when using other than the [default US region](/overview/api/#stacks-and-endpoints), you need to provide the Storage API address: + +```bash +docker run --volume=$("pwd"):/data quay.io/keboola/storage-api-cli:latest create-table in.c-main new-table /data/new-table.csv --token=storage_token --url="https://connection.eu-central-1.keboola.com/" +``` + +Any of the above commands will import the contents of `new-table.csv` in the current directory into the newly +created table. You should see an output similar to this one: + + Authorized as: ondrej.popelka@keboola.com (Odinuv Sandbox) + Bucket found ok + Table create start + Table create end + Table id: in.c-main.new-table + +*Please note that the Docker container can only access folders within the container, so you need to mount a local folder. +In the example above, the local folder `$("pwd")` (replaced by the absolute path at runtime) is mounted as `/data` into the container. +The table is then accessible in this folder. The same approach applies to all other commands working with local files.* + +### Example --- Importing Data +If you only want to import new data into the table, use the `write-table` command and provide +the ID (*bucketName.tableName*) of an existing table. + +To import data into the `new-table` table in the `in.c-main` bucket, use + +```bash +docker run --volume=$("pwd"):/data quay.io/keboola/storage-api-cli:latest write-table in.c-main.new-table /data/new-data.csv --token=storage_token --incremental +``` + +The above command will import the contents of the `new-data.csv` file into the existing table. If the +`--incremental` parameter is supplied, the table contents will be appended. If the parameter is not +supplied, the table contents will be overwritten. You should see an output similar to this one: + + Authorized as: ondrej.popelka@keboola.com (Tutorial) + Table found ok + Import start + Import done in 17 secs. + + Results: + transaction: + warnings: + importedColumns: + - id + - secondCol + totalRowsCount: 8 + totalDataSizeBytes: 4096 + +### Example --- Exporting Data +If you want to export a table from Storage, use the `export-table` command. Provide +the ID (*bucketName.tableName*) of an existing table. + +To export data from the `old-table` table in the `in.c-main` bucket, use + +```bash +docker run --volume=$("pwd"):/data quay.io/keboola/storage-api-cli:latest export-table in.c-main.old-table /data/old-data.csv --token=storage_token +``` + +The above command will export the table from Storage and save it as `old-data.csv` in +the current directory. You should see an output similar to this one: + + Authorized as: ondrej.popelka@keboola.com (Tutorial) + Table found ok + Export done in 17 secs. diff --git a/src/content/docs/storage/api/clients/php-client/index.md b/src/content/docs/storage/api/clients/php-client/index.md new file mode 100644 index 000000000..fb0f80dc6 --- /dev/null +++ b/src/content/docs/storage/api/clients/php-client/index.md @@ -0,0 +1,153 @@ +--- +title: Storage PHP Client Library +slug: 'storage/api/clients/php-client' +redirect_from: + - /integrate/storage/php-client/ +--- + + +The Storage API PHP client library is a portable command line client providing +the most complete [Storage API](https://api.keboola.com/?service=storage) implementation. +It runs on any platform which has PHP installed. +Currently this client implements almost all Storage API functions including, of course, exporting and importing tables. + +The client source is available in our [Github repository](https://github.com/keboola/storage-api-php-client). + +## Installation + +The Library is available as a [Composer package](https://getcomposer.org/). +Unless you already have it, [install Composer](https://getcomposer.org/download/) on your system. +On *nix system, do so by running + +```bash +curl -s https://getcomposer.org/installer | php +mv ./composer.phar ~/bin/composer # or /usr/local/bin/composer +``` + +On Windows, use the [installer](https://getcomposer.org/Composer-Setup.exe). + +To install the library, run + +```bash +composer require keboola/storage-api-client +``` + +in the root of your project. You should get an output similar to this one: + + Using version ^18.10 for keboola/storage-api-client + ./composer.json has been created + Loading composer repositories with package information + Updating dependencies (including require-dev) + - Installing aws/aws-sdk-php (3.18.18) + Downloading: 100% + ... + - Installing keboola/storage-api-client (18.10.0) + Downloading: 100% + Writing lock file + Generating autoload files + +Then add the generated autoloader in your bootstrap script: + +```php +require 'vendor/autoload.php'; +``` + +You can read more in the [Composer documentation](https://getcomposer.org/doc/01-basic-usage.md). Packages +installable by Composer can be browsed at [Packagist package repository](https://packagist.org/). + +## Usage +The Storage API client is implemented as a single class. To create an instance of the class, provide a Storage API token to the +constructor. + +```php + 'your-token', + 'url' => 'https://connection.keboola.com', +]); +``` + +### Example --- Create a Table +To create a new table in Storage, it is recommended to use an additional +[php-csv](https://github.com/keboola/php-csv) library to work +with CSV files. The library will get installed +automatically with the Storage API client, so you can use it out of the box. +To create a new table and import CSV data in it, use the following PHP script: + +```php + 'your-token', + 'url' => 'https://connection.keboola.com', +]); +$csvFile = new CsvFile('./new-table.csv'); +$client->createTableAsync('in.c-main', 'new-table', $csvFile); +``` + +### Example --- Import Data +To import CSV data into an existing table and overwrite its contents, use the following PHP script: + +```php + 'your-token', + 'url' => 'https://connection.keboola.com', +]); +$csvFile = new CsvFile('./new-table.csv'); +$client->writeTableAsync('in.c-main.new-table', $csvFile); +``` + +### Example --- Import Data Incrementally +To import CSV data into an existing table and append the new data to the existing table contents, use the following PHP script: + +```php + 'your-token', + 'url' => 'https://connection.keboola.com', +]); +$csvFile = new CsvFile('./new-table.csv'); +$client->writeTableAsync('in.c-main.new-table', $csvFile, ['incremental' => true]); +``` + +All available upload options are listed in the [API documentation](https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/tables/-id-/import-async). + +### Example --- Export Data +To export data from a Storage table to a CSV file, use the +`TableExporter` class. It is part of the client library. You can use the following script: + +```php + 'your-token', + 'url' => 'https://connection.keboola.com', +]); + +$exporter = new TableExporter($client); +$exporter->exportTable('in.c-main.my-table', './old-table.csv'); +``` diff --git a/src/content/docs/storage/api/clients/python-client/index.md b/src/content/docs/storage/api/clients/python-client/index.md new file mode 100644 index 000000000..54a86a083 --- /dev/null +++ b/src/content/docs/storage/api/clients/python-client/index.md @@ -0,0 +1,121 @@ +--- +title: Python Client Library +slug: 'storage/api/clients/python-client' +redirect_from: + - /integrate/storage/python-client/ +--- + + +The Python client library is a [Storage API client](https://api.keboola.com/?service=storage) which you can use in your Python code. +The current implementation supports all basic data manipulations: + +- Importing data +- Exporting data +- Creating and deleting buckets and tables +- Creating and deleting workspaces + +The client source code is available in our [Github repository](https://github.com/keboola/sapi-python-client/). + +## Installation +The library is published on [PyPI](https://pypi.org/project/kbcstorage/) as `kbcstorage`, so install it with `pip`: + + pip install kbcstorage + +Alternatively, install the latest development version directly from [GitHub](https://github.com/keboola/sapi-python-client): + + pip3 install git+https://github.com/keboola/sapi-python-client.git + +## Usage +The client contains a `Client` class, which encapsulates all API endpoints and holds a storage token and URL. Each API endpoint is +represented by its own class (`Files`, `Buckets`, `Jobs`, etc.), which can be used standalone if you only work with one endpoint. +This means that the two following examples are equivalent: + +```python +from kbcstorage.client import Client + +client = Client('https://connection.keboola.com', 'your-token') +client.tables.detail('in.c-demo.some-table') +``` + +```python +from kbcstorage.tables import Tables + +tables = Tables('https://connection.keboola.com', 'your-token') +tables.detail('in.c-demo.some-table') +``` + +### Example --- Create Table and Import Data +To create a new table in Storage, use the `create` function of the `Tables` class. Provide the name of an existing bucket, +the name of the new table and a CSV file with the table's contents. + +To create the `new-table` table in the `in.c-main` bucket, use: + +```python +from kbcstorage.client import Client + +client = Client('https://connection.keboola.com', 'your-token') +client.tables.create(name='new-table', + bucket_id='in.c-main', + file_path='coords.csv', + primary_key=['id']) +``` + +The above command will import the contents of the `coords.csv` file into the newly created table. It will +also mark the `id` column as the primary key. +### Example --- Load to existing table, incrementally + +To load data incrementally into an existing table, we can use the [load](https://github.com/keboola/sapi-python-client/blob/5a93926c2191ccd6b7402c9e24d9912884d87d4c/kbcstorage/tables.py#L207) method, where `table_id` is the ID of the table that you want to load into, and `path` is the path to your csv file containing the data: + +```python + +from kbcstorage.client import Client + +client = Client('https://connection.keboola.com', 'your-token') + +client.tables.load(table_id=table_id, file_path=path, is_incremental=True) + +``` +### Example --- Export Data +To export data from the `old-table` table in the `in.c-main` bucket, use: + +```python +from kbcstorage.client import Client +import csv + +client = Client('https://connection.keboola.com', 'your-token') +client.tables.export_to_file(table_id='in.c-main.new-table', path_name='.') +with open('./new-table', mode='rt', encoding='utf-8') as in_file: + lazy_lines = (line.replace('\0', '') for line in in_file) + reader = csv.reader(lazy_lines, lineterminator='\n') + for row in reader: + print(row) +``` + +The above command will export the table from Storage into the file `new-table` and read it using +[CSV Reader](https://docs.python.org/3.6/library/csv.html#reader-objects). + +### Other Examples + +```python +# create a client +client = Client('https://connection.keboola.com', 'your-token') + +# create a bucket +client.buckets.create(name='demo', stage='in') + +# list buckets +client.buckets.list() + +# list all tables +client.tables.list() + +# list all tables in a bucket +client.buckets.list_tables(bucket_id='in.c-demo') + +# delete a table +client.tables.delete(table_id='in.c-demo.some-table') + +# delete a bucket +client.buckets.delete(bucket_id='in.c-main', force=True) + +``` diff --git a/src/content/docs/storage/api/clients/r-client/index.md b/src/content/docs/storage/api/clients/r-client/index.md new file mode 100644 index 000000000..226721d36 --- /dev/null +++ b/src/content/docs/storage/api/clients/r-client/index.md @@ -0,0 +1,131 @@ +--- +title: R Client Library +slug: 'storage/api/clients/r-client' +redirect_from: + - /integrate/storage/r-client/ +--- + + +:::note[Limited maintenance] +The R client library is in limited maintenance (last release 2023). For actively maintained access to Storage, prefer the [Python client](/storage/api/clients/python-client/) or the [PHP client](/storage/api/clients/php-client/), or call the [Storage API](https://api.keboola.com/?service=storage) directly. +::: + +The R client library is a [Storage API client](https://api.keboola.com/?service=storage) which you can use in your R code. +The current implementation supports all basic data manipulations: + +- Importing data +- Exporting data +- Creating and deleting buckets and tables + +The client source code is available in our [Github repository](https://github.com/keboola/sapi-r-client). + +## Installation +This library is available on [Github](https://github.com/keboola/sapi-r-client), so we +recommend that you use the `devtools` package to install it. + +```r +# first install the devtools package if it isn't already installed +install.packages("devtools") + +# install dependencies (another github package for aws requests) +devtools::install_github("cloudyr/aws.s3") + +# install the SAPI R client package +devtools::install_github("keboola/sapi-r-client") + +# load the library (dependencies will be loaded automatically) +library(keboola.sapi.r.client) +``` + +## Usage +To list available commands, run +```r +?keboola.sapi.r.client::SapiClient +``` + +**Important**: If you are running the code in R Studio, it might require a restart so that its help index is updated +and the above command works. + +The client is implemented as an [RC class](http://adv-r.had.co.nz/R5.html). To work with it, create an instance of the client. +It requires a valid Storage API token and your stack's Storage API URL (`https://connection.keboola.com` on AWS US; use your [stack's endpoint](/overview/#stacks) otherwise). + +```r +client <- SapiClient$new( + token = 'your-token', + url = 'https://connection.keboola.com' +) +``` + +### Example --- Create a Table and Import Data +To create a new table in Storage, use the `saveTable` function. Provide the name of an existing bucket, +the name of the new table and a CSV file with the table's contents. + +To create the `new-table` table in the `in.c-main` bucket, use + +```r +myDataFrame <- data.frame(id = c(1,2,3,4), secondCol = c('a', 'b', 'c', 'd')) +client <- SapiClient$new( + token = 'your-token', + url = 'https://connection.keboola.com' +) + +table <- client$saveTable( + df = myDataFrame, + bucket = "in.c-main", + tableName = "new-table", + options = list(primaryKey = 'id') +) +``` + +The above command will import the contents of the `myDataFrame` variable into the newly created table. It will +also mark the `id` column as the primary key. + +### Example --- Export Data +If you want to export a table from Storage and import it into R, use the `importTable` function. Provide +the ID (*bucketName.tableName*) of an existing table. + +To export data from the `old-table` table in the `in.c-main` bucket, use + +```r +client <- SapiClient$new( + token = 'your-token', + url = 'https://connection.keboola.com' +) + +data <- client$importTable('in.c-main.old-table') +``` + +The above command will export the table from Storage and save it in the `data` variable. The output is +a [data.table](https://cran.r-project.org/web/packages/data.table/index.html) object compatible with a `data.frame`. + +### Other Examples + +```r +# create a client +client <- SapiClient$new( + token = 'your-token', + url = 'https://connection.keboola.com' +) + +# verify the token +tokenDetails <- client$verifyToken() + +# create a bucket +bucket <- client$createBucket("new_bucket", "in", "A brand new Bucket!") + +# list buckets +buckets <- client$listBuckets() + +# list all tables +tables <- client$listTables() + +# list all tables in a bucket +tables <- client$listTables(bucket = bucket$id) + +# delete a table +client$deleteTable(table$id) + +# delete a bucket +client$deleteBucket(bucket$id) + +``` diff --git a/src/content/docs/storage/api/configurations/index.md b/src/content/docs/storage/api/configurations/index.md new file mode 100644 index 000000000..4e86988e6 --- /dev/null +++ b/src/content/docs/storage/api/configurations/index.md @@ -0,0 +1,525 @@ +--- +title: Component Configurations API +slug: 'storage/api/configurations' +redirect_from: + - /integrate/storage/api/configurations/ +--- + + +[Configurations](/storage/configurations/) are an important part of a Keboola project. Most operations are +available in the UI. Use the API if you want to manipulate the configurations programmatically. + +Configurations represent component **instances** in a project. Each Keboola component has different configuration +options and requirements, which must be respected. As such, Keboola configurations provide a general framework for configuring +components, while the specific implementation details are left to the components themselves. + +When working with the [Component Configurations API](https://api.keboola.com/?service=storage#tag--Component-Configurations), +you need to know the `componentId` of the component being configured. +You can see a list of public components in [the Developer Portal](https://components.keboola.com/components), or you can get +a list of all available components with the [API index call](https://api.keboola.com/?service=storage#get-/v2/storage). +See our [example](https://documenter.getpostman.com/view/3086797/kbc-samples/77h845D?version=latest#9b9f3e7b-de3b-4c90-bad6-a8760e3852eb). + +It will give you something like this: + +```json +{ + "host": "4edece0b0052", + "api": "storage", + "version": "v2", + "revision": "21fb56a0f6d61a307f350247a45950b1e4049625", + "documentation": "https://connection.keboola.com/api/storage/doc.json", + "components": [ + { + "id": "keboola.ex-aws-s3", + "type": "extractor", + "name": "AWS S3", + "description": "AWS Simple Storage Service", + "longDescription": "Download ... from AWS S3 and upload them to Storage.", + "version": 23, + "hasUI": false, + "hasRun": false, + "ico32": "https://ui.keboola-assets.com/.../keboola.ex-aws-s3/32/20.png", + "ico64": "https://ui.keboola-assets.com/.../keboola.ex-aws-s3/64/20.png", + "data": { + "definition": { + "type": "aws-ecr", + "uri": "147946154733.../keboola.ex-aws-s3", + "tag": "v3.0.0", + "repository": { + "region": "us-east-1" + } + }, + "vendor": { + "contact": [ + "Keboola", + "Křižíkova 488/115\n186 00 Prague 8\nCzech Republic", + "support@keboola.com" + ], + "licenseUrl": "https://github.com/keboola/aws-s3-extractor/blob/master/LICENSE" + }, + "configuration_format": "json", + "network": "bridge", + "memory": "512m", + "forward_token": false, + "forward_token_details": false, + "default_bucket": true, + "default_bucket_stage": "in", + "staging_storage": { + "input": "local" + } + }, + "flags": [ + "genericDockerUI", + "genericDockerUI-processors", + "appInfo.dataIn" + ], + "configurationSchema": {}, + "emptyConfiguration": {}, + "uiOptions": {}, + "configurationDescription": null, + "documentationUrl": "https://help.keboola.com/extractors/other/aws-s3/" + } + ], + "services": [...], + "urlTemplates": {...} +} +``` + +From here, you can see all available information about a particular component. In the following examples, we +will use `keboola.ex-aws-s3` --- the AWS S3 extractor. + +## Configuration Structure +Component configurations are largely dependent on the actual component being configured. This makes creating configurations manually +a bit tricky. Rather than starting from scratch, we recommend creating a configuration through the UI and then modifying it when you understand it. + +### Inspecting Configuration +To obtain an existing configuration, you can either use the list of configurations above or +the [Configuration Detail](https://api.keboola.com/?service=storage#get-/v2/storage/branch/-branchId-/components/-componentId-/configs/-configurationId-) +API call. See an [example](https://documenter.getpostman.com/view/3086797/kbc-samples/77h845D?version=latest#9b9f3e7b-de3b-4c90-bad6-a8760e3852eb) for obtaining a +configuration of the `keboola.ex-aws-s3` component. You will receive a response similar to this: + +```json +{ + "id": "364479526", + "name": "test", + "description": "", + "created": "2018-03-08T14:54:19+0100", + "creatorToken": { + "id": 27865, + "description": "ondrej.popelka@keboola.com" + }, + "version": 5, + "changeDescription": "Table first table edited", + "isDeleted": false, + "configuration": { + "parameters": { + "accessKeyId": "AKIAIBZYEEXQILP46FCA", + "#secretAccessKey": "KBC::ComponentProjectEncrypted==p5gvUw4RSGiVJjT2ayVORpqS7yiKhExi7NnQECntVm8haHaHtFNVDMT8X8b+htnixpXhPIQ9yV+ETrvr+hNeYfh+Ex+UpC//QPWnLcEOC8XOLgmQN8BNgRGSERWUziK0" + } + }, + "rowsSortOrder": [], + "rows": [ + { + "id": "364481153", + "name": "first table", + "description": "", + "configuration": { + "parameters": { + "bucket": "travis-php-db-import-tests-s3filesbucket-vm9zhtm5jd7s", + "key": "tw_accounts.csv", + "saveAs": "first-table", + "includeSubfolders": false, + "newFilesOnly": true + }, + "processors": { + "after": [ + { + "definition": { + "component": "keboola.processor-move-files" + }, + "parameters": { + "direction": "tables", + "addCsvSuffix": true + } + }, + { + "definition": { + "component": "keboola.processor-create-manifest" + }, + "parameters": { + "delimiter": ",", + "enclosure": "\"", + "incremental": false, + "primary_key": [], + "columns": [], + "columns_from": "header" + } + }, + { + "definition": { + "component": "keboola.processor-skip-lines" + }, + "parameters": { + "lines": 1 + } + } + ] + } + }, + "isDisabled": false, + "version": 3, + "created": "2018-03-08T14:58:33+0100", + "creatorToken": { + "id": 27865, + "description": "ondrej.popelka@keboola.com" + }, + "changeDescription": "Table first table edited", + "state": { + "lastDownloadedFileTimestamp": "1511176959", + "processedFilesInLastTimestampSecond": [ + "tw_accounts.csv" + ] + } + } + ], + "state": {}, + "currentVersion": { + "created": "2018-03-08T23:27:37+0100", + "creatorToken": { + "id": 27865, + "description": "ondrej.popelka@keboola.com" + }, + "changeDescription": "Table first table edited" + } +} +``` + +The actual component configuration is split into three parts: + +- `configuration` node, containing an arbitrary component configuration +- `state` node, containing a component [state file](/extend/common-interface/config-file/#state-file) +- `rows` node, containing iterations of `configuration` and `state` + +The important part is the ID of the configuration you want to work with. In the following examples, we will use +`364479526`. + +### Configuration +The `configuration` node maps to the [configuration file](/extend/common-interface/config-file/#configuration-file-structure). +It can contain the `storage`, `parameters`, `processors` and `authorization` child nodes (the `image_parameters` and `action` nodes found in the config file +are injected at runtime and are not stored in the configuration). The `authorization` node is set in the configuration only when +[credentials injection](/extend/common-interface/oauth/#credentials-injection) should be used, otherwise it is also set during the runtime. +The `processors` node defines the [processors and their configuration](/extend/component/processors/). +The most common sub-nodes stored in the `configuration` node are therefore `parameters` (containing an arbitrary component configuration) +and `storage` (containing [input](/extend/component/tutorial/input-mapping/) and [output mapping](/extend/component/tutorial/output-mapping/)). +Both are transferred to the +configuration file without modification; that means that the [`storage` configuration](/extend/common-interface/config-file/#configuration-file-structure) +is directly usable in the `configuration` node. The `parameters` node is fully dependent on the component and has no universal specification or rules. + +In the above example, the `configuration` node contains the following: + +```json +"parameters": { + "accessKeyId": "AKIAIBZYEEXQILP46FCA", + "#secretAccessKey": "KBC::ComponentProjectEncrypted==p5gvUw4RSGiVJjT2ayVORpqS7yiKhExi7NnQECntVm8haHaHtFNVDMT8X8b+htnixpXhPIQ9yV+ETrvr+hNeYfh+Ex+UpC//QPWnLcEOC8XOLgmQN8BNgRGSERWUziK0" +} +``` + +That means that the component is not using input mapping nor output mapping. The allowed contents of `parameters` are described +in the [AWS S3 extractor code documentation](https://github.com/keboola/aws-s3-extractor#configuration-options). + +### Configuration Rows +The `rows` node contains iterations of the configuration. The interpretation of configuration rows is again dependent on the +component implementation. In the presented case of the `keboola.ex-aws-s3` component, each row corresponds to a single extracted table. +When `rows` node is non-empty, the component behavior is slightly modified. It behaves as if it were executed as many times as +there are rows. For each row, the `configuration` node from `root` and the `configuration` node from `rows` are merged, with +the latter overwriting the former in the case of conflict. + +Given the above configuration, the **effective configuration** passed to the component +[configuration file](/extend/common-interface/config-file/#configuration-file-structure) will be as follows: + +```json +{ + "parameters": { + "accessKeyId": "AKIAIBZYEEXQILP46FCA", + "#secretAccessKey": "KBC::ComponentProjectEncrypted==p5gvUw4RSGiVJjT2ayVORpqS7yiKhExi7NnQECntVm8haHaHtFNVDMT8X8b+htnixpXhPIQ9yV+ETrvr+hNeYfh+Ex+UpC//QPWnLcEOC8XOLgmQN8BNgRGSERWUziK0", + "bucket": "travis-php-db-import-tests-s3filesbucket-vm9zhtm5jd7s", + "key": "tw_accounts.csv", + "saveAs": "first-table", + "includeSubfolders": false, + "newFilesOnly": true + } +} +``` + +The first two parameters (`accessKeyId` and `#secretAccessKey`) are taken from the root `configuration`, the other +parameters are taken from the first rows' `configuration`. The `processors` node is never passed to the configuration file. +With the above configuration, the component will be executed only once, because there is one row. If there are no rows, the +component will still be executed once. If there were two rows, the component would be executed twice. + +If the component is executed more than once, the operations are executed in the following order: + +- input mapping for the first row +- run with the first row configuration (merged with root configuration) +- output mapping for the first row +- input mapping for the second row +- run with the second row configuration (merged with root configuration) +- output mapping for the second row + +All of these are executed in a single [job](/integrate/jobs/). However, even though multiple rows are executed in a single +job, the actual executions are still completely isolated. I.e., there is no way to share anything between the rows +(apart from the common `configuration`). It also means that the outputs of the first row are available in the Keboola project before +the second row starts, and the inputs for the second row are read only after the first row finishes processing. + +What is considered 'first' and 'second' -- i.e. the order of rows -- is defined by the order of items in the `rows` array. +See [below](#modifying-a-configuration) for an example of modifying the row order. + +Theoretically, configuration rows are supported for every component as long as the effective configuration matches what +the component expects. Configuration rows can be used to split the configuration into a common part (typically credentials) and an +iterable part which is repeated many times. Keep in mind that configurations heavily modified through the API might **not be supported +in the UI**. + +### State +The `state` node contains the content of the [state file](/extend/common-interface/config-file/#state-file). The +`state` is read from the state file and then supplied to the state file on the next run. In the above configuration, +the state is: + +```json +{ + "lastDownloadedFileTimestamp": "1511176959", + "processedFilesInLastTimestampSecond": [ + "tw_accounts.csv" + ] +} +``` + +`State` is considered an internal property of a component and you should avoid modifying it. The only reasonable modification of +`state` is to delete it -- in that case, the configuration will run as if it were run for the first time. To delete the `state`, set it to `{}`. +If configuration rows are used, then the `state` is stored separately for each row and the `state` node in configuration root is +not used. + +## Working with Configurations +Here, the most common operations done with configurations are described in examples. Feel free to go through the +[API reference](https://api.keboola.com/?service=storage#tag--Component-Configurations) for a full authoritative list of configuration features. + +### List Configurations +To obtain configuration details, use the [List Configs call](https://api.keboola.com/?service=storage#get-/v2/storage/branch/-branchId-/components/-componentId-/configs), +which will return all the configuration details. This means + +- the configuration itself (`configuration`) --- [section on configuration](#modifying-a-configuration) follows; +- configuration rows (`rows`) --- additional data of the configuration; and +- configuration state (`state`) --- [component state](/extend/common-interface/config-file/#state-file). + +Please note that the contents +of the `configuration`, `rows` and `state` sections depend purely on the component itself. See an [example](https://documenter.getpostman.com/view/3086797/kbc-samples/77h845D?version=latest#9b9f3e7b-de3b-4c90-bad6-a8760e3852eb). + +A sample result for the AWS S3 extractor looks like this: + +```json +[ + { + "id": "364479526", + "name": "test", + "description": "", + "created": "2018-03-08T14:54:19+0100", + "creatorToken": { + "id": 27865, + "description": "ondrej.popelka@keboola.com" + }, + "version": 4, + "changeDescription": "Table first table edited", + "isDeleted": false, + "configuration": { + "parameters": { + "accessKeyId": "AKIAIBZYEEXQILP46FCA", + "#secretAccessKey": "KBC::ComponentProjectEncrypted==p5gvUw4RSGiVJjT2ayVORpqS7yiKhExi7NnQECntVm8haHaHtFNVDMT8X8b+htnixpXhPIQ9yV+ETrvr+hNeYfh+Ex+UpC//QPWnLcEOC8XOLgmQN8BNgRGSERWUziK0" + } + }, + "rowsSortOrder": [], + "rows": [ + { + "id": "364481153", + "name": "first table", + "description": "", + "configuration": {...}, + "isDisabled": false, + "version": 2, + "created": "2018-03-08T14:58:33+0100", + "creatorToken": { + "id": 27865, + "description": "ondrej.popelka@keboola.com" + }, + "changeDescription": "Table first table edited", + "state": {} + } + ], + "state": {}, + "currentVersion": { + "created": "2018-03-08T15:21:28+0100", + "creatorToken": { + "id": 27865, + "description": "ondrej.popelka@keboola.com" + }, + "changeDescription": "Table first table edited" + } + } +] +``` + +### Modifying Configuration +**Note: Configurations modified through the API might not be editable in the Keboola UI.** They can be run or used in an orchestration without any problems. + +Modifying a configuration means that a new version of that configuration is created. +For modifying a configuration, use the +[Update Configuration](https://api.keboola.com/?service=storage#put-/v2/storage/branch/-branchId-/components/-componentId-/configs/-configurationId-) API call. +See an [example](https://documenter.getpostman.com/view/3086797/kbc-samples/77h845D?version=latest#9b9f3e7b-de3b-4c90-bad6-a8760e3852eb) in which the +configuration is modified to the following to set new credentials: + +```json +{ + "parameters": { + "accessKeyId": "a", + "#secretAccessKey": "b" + } +} +``` + +Notice that the configuration must be sent in the form field `configuration` as the endpoint does not accept pure JSON (yet). +Take great care to pass **only the contents** of the `configuration` node as in the above example. The configuration **must not be wrapped** in the +`configuration` node, otherwise the component will not +receive the configuration it expects. Also take care to properly escape the JSON using [URL encoding](https://en.wikipedia.org/wiki/Percent-encoding), +otherwise it may be misinterpreted. The raw HTTP request should look similar to this: + + curl --request PUT \ + --url https://connection.keboola.com/v2/storage/components/keboola.ex-aws-s3/configs/364479526 \ + --header "Content-Type: application/json" \ + --header 'X-StorageAPI-Token: {{token}}' \ + --data-binary "{ + \"configuration\": { + \"parameters\": { + \"accessKeyId\": \"a\", + \"#secretAccessKey\": \"b\" + } + } + }" + +Also note that the entire configuration must be always sent, there is no way to patch only part of it. +The same way the `configuration` is modified, other properties can be modified too. For example, you may want to +reset `state` by setting it to `{}`, or you can change the order of the configuration rows by setting the `rowsSortOrder` property. +The `rowsSortOrder` is an array of row ids -- see an [example](https://documenter.getpostman.com/view/3086797/kbc-samples/77h845D?version=latest#9b9f3e7b-de3b-4c90-bad6-a8760e3852eb) (Set Row order of S3 extractor) +for the exact example request. + +### Modifying Configuration Row +Very similar to modifying a configuration, modifying a configuration **row** means that a new version of +the **entire configuration** is created. For modifying a configuration row, use the +[Update Row](https://api.keboola.com/?service=storage#put-/v2/storage/branch/-branchId-/components/-componentId-/configs/-configurationId-/rows/-rowId-) API call. + +See an [example](https://documenter.getpostman.com/view/3086797/kbc-samples/77h845D?version=latest#9b9f3e7b-de3b-4c90-bad6-a8760e3852eb) in which the +configuration row is modified to: + +```json +{ + "parameters": { + "bucket": "some-bucket", + "key": "sample.csv", + "includeSubfolders": false, + "newFilesOnly": true + } +} +``` + +The rules for updating a configuration row are the same as for [updating a configuration](#modifying-a-configuration). Also note that +a configuration row is never evaluated alone, it is always merged with the root `configuration`. If the same properties are defined +in the root `configuration` and row `configuration`, the values from the row are used. There is also an +[example](https://documenter.getpostman.com/view/3086797/kbc-samples/77h845D?version=latest#9b9f3e7b-de3b-4c90-bad6-a8760e3852eb) of how to reset the row +state by setting `state` to `{}`. + +### Configuration Versions +When you [update a configuration](https://api.keboola.com/?service=storage#put-/v2/storage/branch/-branchId-/components/-componentId-/configs/-configurationId-), +a new configuration version is actually created. In the above calls, only the last (active/published) configuration +is returned. To obtain a list of all recorded versions, use the +[List Versions API call](https://api.keboola.com/?service=storage#get-/v2/storage/branch/-branchId-/components/-componentId-/configs/-configurationId-/versions). +See this [example](https://documenter.getpostman.com/view/3086797/kbc-samples/77h845D?version=latest#9b9f3e7b-de3b-4c90-bad6-a8760e3852eb) +which would give you an output similar to the one below: + +```json +[ + { + "version": 4, + "created": "2018-03-08T15:21:28+0100", + "creatorToken": { + "id": 27865, + "description": "ondrej.popelka@keboola.com" + }, + "changeDescription": "Table first table edited", + "isDeleted": false, + "name": "test", + "description": "" + }, + { + "version": 3, + "created": "2018-03-08T14:58:33+0100", + "creatorToken": { + "id": 27865, + "description": "ondrej.popelka@keboola.com" + }, + "changeDescription": "Table first table added", + "isDeleted": false, + "name": "test", + "description": "" + }, + { + "version": 2, + "created": "2018-03-08T14:55:50+0100", + "creatorToken": { + "id": 27865, + "description": "ondrej.popelka@keboola.com" + }, + "changeDescription": "AWS Credentials edited", + "isDeleted": false, + "name": "test", + "description": "" + }, + { + "version": 1, + "created": "2018-03-08T14:54:19+0100", + "creatorToken": { + "id": 27865, + "description": "ondrej.popelka@keboola.com" + }, + "changeDescription": "", + "isDeleted": false, + "name": "test", + "description": "" + } +] +``` + +The field `version` represents the `version_id` in the following API example. + +### Rollback Configuration +After choosing a particular version, you can revert to that version by +[rolling back](https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/components/-componentId-/configs/-configurationId-/versions/-versionId-/rollback), +i.e., making a new version identical to the chosen one. See an [example](https://documenter.getpostman.com/view/3086797/kbc-samples/77h845D#2050856a-66b3-4120-9552-d1278a96621e) +of how to rollback the configuration `364479526` of the `keboola.ex-aws-s3` component to version `3`. + +It will create a new version of the configuration and return the ID of the version: +```json +{ + "version": "26" +} +``` + +### Creating Configuration Copy +After choosing a particular version, you can create a new independent +[configuration copy](https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/components/-componentId-/configs/-configurationId-/versions/-versionId-/create) +of it. See an [example](https://documenter.getpostman.com/view/3086797/kbc-samples/77h845D?version=latest#9b9f3e7b-de3b-4c90-bad6-a8760e3852eb) +of how to create a new configuration called `test-copy` from version `3` of the `364479526` configuration +for the `keboola.ex-aws-s3` component. + +It will return the ID of the newly created configuration: +```json +{ + "id": "364494012" +} +``` + diff --git a/src/content/docs/storage/api/import-export/index.md b/src/content/docs/storage/api/import-export/index.md new file mode 100644 index 000000000..8a49d0b09 --- /dev/null +++ b/src/content/docs/storage/api/import-export/index.md @@ -0,0 +1,342 @@ +--- +title: Manually Importing and Exporting Data +slug: 'storage/api/import-export' +redirect_from: + - /integrate/storage/api/import-export/ +--- + + +## Working with Data +Keboola Table Storage (Tables) and Keboola File Storage (File Uploads) are heavily connected together. +Keboola File Storage is technically a layer on top of the Amazon S3 service, and Keboola Table +Storage is a layer on top of a [database backend](/storage/#backends). + +To upload a table, take the following steps: + +- Request a [file upload](https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/files/prepare) from +Keboola File Storage. You will be given a destination for the uploaded file on an S3 server. +- Upload the file there. When the upload is finished, the data file will be available in the *File Uploads* section. +- Initiate an [asynchronous table import](https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/tables/-id-/import-async) +from the uploaded file (use it as the `dataFileId` parameter) into the destination table. +The import is asynchronous, so the request only creates a job and you need to poll for its results. +The imported files must conform to the [RFC4180 Specification](https://tools.ietf.org/html/rfc4180). + +![Schema of file upload process](/storage/api/async-import-handling.svg) + +Exporting a table from Storage is analogous to its importing. First, data is [asynchronously +exported](https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/tables/-id-/export-async) from +Table Storage into File Uploads. Then you can request to [download +the file](https://api.keboola.com/?service=storage#get-/v2/storage/branch/-branchId-/files/-fileId-), which will give you +access to an S3 server for the actual file download. + +### Manually Uploading a File +To upload a file to Keboola File Storage, follow the instructions outlined in the +[API documentation](https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/files/prepare). +First create a file resource; to create a new file called +[`new-file.csv`](/storage/api/new-table.csv) with `52` bytes, call: + +```bash +curl --request POST --header "Content-Type: application/json" --header "X-StorageApi-Token:storage-token" --data-binary "{ \"name\": \"new-file.csv\", \"sizeBytes\": 52, \"federationToken\": 1 }" https://connection.keboola.com/v2/storage/files/prepare +``` + +Which will return a response similar to this: + +```json +{ + "id": 192726698, + "created": "2016-06-22T10:44:35+0200", + "isPublic": false, + "isSliced": false, + "isEncrypted": false, + "name": "new_file2.csv", + "url": "https://s3.amazonaws.com/kbc-sapi-files/exp-15/1134/files/2016/06/22/192726697.new_file2?X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIAJ2N244XSWYVVYVLQ%2F20160622%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20160622T084435Z&X-Amz-SignedHeaders=host&X-Amz-Expires=3600&X-Amz-Signature=86136cced74cdf919953cde9e2a0b837bd0b8f147aa6b7b30c2febde3b92d83d", + "region": "us-east-1", + "sizeBytes": 52, + "tags": [], + "maxAgeDays": 15, + "runId": null, + "runIds": [], + "creatorToken": { + "id": 53044, + "description": "ondrej.popelka@keboola.com" + }, + "uploadParams": { + "key": "exp-15/1134/files/2016/06/22/192726697.new_file2.csv", + "bucket": "kbc-sapi-files", + "acl": "private", + "credentials": { + "AccessKeyId": "ASI...H7Q", + "SecretAccessKey": "QbO...7qu", + "SessionToken": "Ago...bsF", + "Expiration": "2016-06-22T20:44:35+00:00" + } + } +} +``` + +The important parts are: `id` of the file, which will be needed later, the `uploadParams.credentials` node, +which gives you credentials to AWS S3 to upload your file, and +the `key` and `bucket` nodes, which define the target S3 destination as *s3://`bucket`/`key`*. +To upload the files to S3, you need an S3 client. There are a large number of clients available: +for example, use the +[S3 AWS command line client](https://docs.aws.amazon.com/cli/latest/userguide/cli-chap-install.html). +Before using it, [pass the credentials](https://docs.aws.amazon.com/cli/latest/topic/config-vars.html#credentials) +by executing, for instance, the following commands + +on *nix systems: +```bash +export AWS_ACCESS_KEY_ID=ASI...H7Q +export AWS_SECRET_ACCESS_KEY=QbO...7qu +export AWS_SESSION_TOKEN=Ago...wU= +``` + +or on Windows: +```bash +SET AWS_ACCESS_KEY_ID=ASI...H7Q +SET AWS_SECRET_ACCESS_KEY=QbO...7qu +SET AWS_SESSION_TOKEN=Ago...bsF +``` + +Then you can actually upload the `new-table.csv` file by executing the AWS S3 CLI [cp command](https://docs.aws.amazon.com/cli/latest/reference/s3/cp.html): +```bash +aws s3 cp new-table.csv s3://kbc-sapi-files/exp-15/1134/files/2016/06/22/192726697.new_file2.csv +``` + +After that, import the file into Table Storage, by calling either +[Create Table API call](https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/buckets/-id-/tables-async) +(for a new table) or +[Load Data API call](https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/tables/-id-/import-async) +(for an existing table). + +```bash +curl --request POST --header "Content-Type: application/json" --header "X-StorageApi-Token:storage-token" --data-binary "{ \"dataFileId\": 192726698, \"name\": \"new-table\" }" https://connection.keboola.com/v2/storage/buckets/in.c-main/tables-async +``` + +This will create an asynchronous job, importing data from the `192726698` file into the `new-table` destination table in the `in.c-main` bucket. +Then [poll for the job results](/integrate/jobs/#job-polling), or review its status in the UI. + +#### Python Example +The above process is implemented in the following example script in Python. This script uses the +[Requests](https://2.python-requests.org/en/master/) library for sending HTTP requests and +the [Boto 3](https://github.com/boto/boto3) library for working with Amazon S3. Both libraries can be +installed using pip: + +```bash +pip install boto3 +pip install requests +``` + +```python +import requests +import os +import json +import boto3 +from time import sleep + +storageToken = 'yourToken' +# Source filename (including path) +fileName = 'simple.csv' +# Target Storage Bucket (assumed to exist) +bucketName = 'in.c-main' +# Target Storage Table (assumed NOT to exist) +tableName = 'my-new-table' + +print('\nCreating upload file') + +# Create a new file in Storage +# See https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/files/prepare +response = requests.post( + 'https://connection.keboola.com/v2/storage/files/prepare', + data={ + 'name': fileName, + 'sizeBytes': os.stat(fileName).st_size, + 'federationToken': 1 + }, + headers={'X-StorageApi-Token': storageToken} +) +parsed = json.loads(response.content.decode('utf-8')) +# print(response.request.body) +# print(json.dumps(parsed, indent=4)) + +# Get AWS Credentials +accessKeyId = parsed['uploadParams']['credentials']['AccessKeyId'] +accessKeySecret = parsed['uploadParams']['credentials']['SecretAccessKey'] +sessionToken = parsed['uploadParams']['credentials']['SessionToken'] +region = parsed['region'] +fileId = parsed['id'] + +print('\nUploading to S3') + +# Upload file to S3 +# See https://boto3.amazonaws.com/v1/documentation/api/latest/guide/configuration.html +s3 = boto3.resource('s3', region_name=region, aws_access_key_id=accessKeyId, aws_secret_access_key=accessKeySecret, aws_session_token=sessionToken) +data = open(fileName, 'rb') +s3.Bucket(parsed['uploadParams']['bucket']).put_object(Key=parsed['uploadParams']['key'], Body=data) + +print('\nCreating table') + +# Load data from file into the Storage table +# See https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/buckets/-id-/tables-async +response = requests.post( + 'https://connection.keboola.com/v2/storage/buckets/%s/tables-async' % bucketName, + data={'name': tableName, 'dataFileId': fileId, 'delimiter': ',', 'enclosure': '"'}, + headers={'X-StorageApi-Token': storageToken}, +) +parsed = json.loads(response.content.decode('utf-8')) +# print(json.dumps(parsed, indent=4)) +if (parsed['status'] == 'error'): + print(parsed['error']) + exit(2) + +status = parsed['status'] +while (status == 'waiting') or (status == 'processing'): + print('\nWaiting for import to finish') + # See https://api.keboola.com/?service=storage#get-/v2/storage/jobs/-jobId- + response = requests.get(parsed['url'], headers={'X-StorageApi-Token': storageToken}) + jobParsed = json.loads(response.content.decode('utf-8')) + status = jobParsed['status'] + sleep(1) + +# print(json.dumps(jobParsed, indent=4)) +if (jobParsed['status'] == 'error'): + print(jobParsed['error']['message']) + exit(2) +``` + +#### Upload Files Using Storage API Importer +For production setup, we recommend using the approach [outlined above](#manually-uploading-a-file) +with direct upload to S3 as it is more reliable and universal. +In case you need to avoid using an S3 client, it is also possible to upload the +file by a simple HTTP request to [Storage API Importer Service](/storage/api/importer/). + +```bash +curl --request POST --header "X-StorageApi-Token:storage-token" --form "data=@new-file.csv" https://import.keboola.com/upload-file +``` + +The above will return a response similar to this: + +```json +{ + "id": 418137780, + "created": "2018-07-17T13:48:57+0200", + "isPublic": false, + "isSliced": false, + "isEncrypted": true, + "name": "404.md", + "url": "https:\/\/kbc-sapi-files.s3.amazonaws.com\/exp-15\/4088\/files\/2018\/07\/17\/418137779.new-file.csv...truncated", + "region": "us-east-1", + "sizeBytes": 1765, + "tags": [], + "maxAgeDays": 15, + "runId": null, + "runIds": [], + "creatorToken": { + "id": 144880, + "description": "file upload" + } +} +``` + +After that, import the file into Table Storage by calling either +[Create Table API call](https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/buckets/-id-/tables-async) +(for a new table) or +[Load Data API call](https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/tables/-id-/import-async) +(for an existing table). + +### Working with Sliced Files +Depending on the backend and table size, the data file may be sliced into chunks. +Requirements for uploading sliced files are described in the respective part of the +[API documentation](https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/files/prepare). + +When you attempt to download a sliced file, you will instead obtain its manifest +listing the individual parts. Download the parts individually and join them +together. For a reference implementation of this process, see +our [TableExporter class](https://github.com/keboola/storage-api-php-client/blob/master/src/Keboola/StorageApi/TableExporter.php). + +**Important:** When exporting a table through the *Table* --- *Export* UI, the file will +be already merged and listed in the *File Uploads* section with the `storage-merged-export` tag. + +If you want to download a sliced file, [get credentials](https://api.keboola.com/?service=storage#get-/v2/storage/branch/-branchId-/files/-fileId-) +to download the file from AWS S3. Assuming that the file ID is 192611596, for example, call + +```bash +curl --header "X-StorageAPI-Token: storage-token" https://connection.keboola.com/v2/storage/files/192611596?federationToken=1 +``` + +which will return a response similar to this: + +```json +{ + "id": 192611596, + "created": "2016-06-21T15:25:35+0200", + "name": "in.c-redshift.blog-data.csv", + "url": "https://s3.amazonaws.com/kbc-sapi-files/exp-2/578/table-exports/in/c-redshift/blog-data/192611594.csvmanifest?X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Credential=AKIAJ2N244XSWYVVYVLQ%2F20160621%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20160621T135137Z&X-Amz-SignedHeaders=host&X-Amz-Expires=3600&X-Amz-Signature=ee69d94f0af06bcf924df0f710dcd92e6503a13c8a11a86be2606552bf9a8b26", + "region": "us-east-1", + "sizeBytes": 24541, + "tags": [ + "table-export" + ], + ... + "s3Path": { + "bucket": "kbc-sapi-files", + "key": "exp-2/578/table-exports/in/c-redshift/blog-data/192611594.csv" + }, + "credentials": { + "AccessKeyId": "ASI...UQQ", + "SecretAccessKey": "LHU...HAp", + "SessionToken": "Ago...uwU=", + "Expiration": "2016-06-22T01:51:37+00:00" + } +} +``` + +The field `url` contains the URL to the file manifest. Upon downloading it, you will get a JSON file with contents +similar to this: + +```json +{ + "entries": [ + {"url":"s3://kbc-sapi-files/exp-2/578/table-exports/in/c-redshift/blog-data/192611594.csv0000_part_00"}, + {"url":"s3://kbc-sapi-files/exp-2/578/table-exports/in/c-redshift/blog-data/192611594.csv0001_part_00"} + ] +} +``` + +Now you can download the actual data file slices. URLs are provided in the manifest file, and credentials to them +are returned as part of the previous file info call. To download the files from S3, you need an S3 client. There +are a wide number of clients available; for example, use the +[S3 AWS command line client](https://docs.aws.amazon.com/cli/latest/userguide/cli-chap-install.html). Before +using it, [pass the credentials](https://docs.aws.amazon.com/cli/latest/topic/config-vars.html#credentials) +by executing , for instance, the following commands + +on *nix systems: +```bash +export AWS_ACCESS_KEY_ID=ASI...UQQ +export AWS_SECRET_ACCESS_KEY=LHU...HAp +export AWS_SESSION_TOKEN=Ago...wU= +``` + +or on Windows: +```bash +SET AWS_ACCESS_KEY_ID=ASI...UQQ +SET AWS_SECRET_ACCESS_KEY=LHU...HAp +SET AWS_SESSION_TOKEN=Ago...wU= +``` + +Then you can actually download the files by executing the AWS S3 CLI [cp command](https://docs.aws.amazon.com/cli/latest/reference/s3/cp.html): +```bash +aws s3 cp s3://kbc-sapi-files/exp-2/578/table-exports/in/c-redshift/blog-data/192611594.csv0000_part_00 192611594.csv0000_part_00 +aws s3 cp s3://kbc-sapi-files/exp-2/578/table-exports/in/c-redshift/blog-data/192611594.csv0001_part_00 192611594.csv0001_part_00 +``` + +After that, merge the files together by executing the following commands + +on *nix systems: +```bash +cat 192611594.csv0000_part_00 192611594.csv0001_part_00 > merged.csv +``` + +or on Windows: +```bash +copy 192611594.csv0000_part_00 /B +192611594.csv0001_part_00 /B merged2.csv +``` diff --git a/src/content/docs/storage/api/importer/index.md b/src/content/docs/storage/api/importer/index.md new file mode 100644 index 000000000..b60099b60 --- /dev/null +++ b/src/content/docs/storage/api/importer/index.md @@ -0,0 +1,50 @@ +--- +title: Storage API Importer +slug: 'storage/api/importer' +redirect_from: + - /integrate/storage/api/importer/ +--- + + +The [whole process of importing](/storage/api/) a table into Storage can be simplified with the +Storage API Importer Service. +The Storage API Importer allows you to make an HTTP POST request and import a file directly into an existing Storage table. + +The HTTP request must contain the `tableId` and `data` form fields. The specified table must already exist in [Storage](/storage/). +Therefore to upload the `my-table.csv` CSV file (and replace the contents) into the `my-table` table in the `in.c-main` bucket, +call: + +```bash +curl --request POST --header "X-StorageApi-Token:storage-token" --form "tableId=in.c-main.my-table" --form "data=@my-table.csv" "https://import.keboola.com/write-table" +``` + +Using the Storage API Importer is the easiest way to upload data into Storage (except for +using one of the [API clients](/storage/api/clients/)). However, the disadvantage is that the whole data file +has to be posted in a single HTTP request. **The maximum limit for a file size is 2GB and the transfer time is 45 minutes**. +This means that for substantially large files (usually more than hundreds of MB) +you may experience timeouts. If that happens, use the above outlined approach and upload the +files [directly to S3](/storage/api/import-export/#manually-uploading-a-file). + +## Parameters + +- `tableId` (required) Storage Table ID, example: in.c-main.users +- `data` (required) Uploaded CSV file. Raw file or compressed by [gzip](http://www.gzip.org/) +- `delimiter` (optional) Field delimiter used in a CSV file. The default value is ' , '. Use '\t' or type the tab char for tabulator. +- `enclosure` (optional) Field enclosure used in a CSV file. The default value is '"'. +- `escapedBy` (optional) CSV escape character; empty by default. +- `incremental` (optional) If incremental is set to 0 (its default), the target table is truncated before each import. + +Full list of available parameters is available in the [API documentation](https://api.keboola.com/?service=import#import). + +## Examples +To load data incrementally (append new data to existing contents): + +```bash +curl --request POST --header "X-StorageApi-Token:storage-token" --form "incremental=1" --form "tableId=in.c-main.my-table" --form "data=@my-table.csv" "https://import.keboola.com/write-table" +``` + +To load data with a non-default delimiter (tabulator) and enclosure (empty): + +```bash +curl --request POST --header "X-StorageApi-Token:storage-token" --form "delimiter=\t" --form "enclosure=" --form "tableId=in.c-main.my-table" --form "data=@my-table.csv" "https://import.keboola.com/write-table" +``` diff --git a/src/content/docs/storage/api/index.md b/src/content/docs/storage/api/index.md new file mode 100644 index 000000000..2f3edc9b1 --- /dev/null +++ b/src/content/docs/storage/api/index.md @@ -0,0 +1,37 @@ +--- +title: Storage API +slug: 'storage/api' +redirect_from: + - /integrate/storage/ + - /integrate/storage/api/ +--- + + +If you are new to Keboola, you should make yourself familiar with +the [Storage component](/storage/) before you start using it. +For a general introduction to working with Keboola APIs, see the [API Introduction](/overview/api/). +[Storage API](https://api.keboola.com/?service=storage) provides a number of functions. These are the most important ones: + +- [Component configurations](https://api.keboola.com/?service=storage#tag--Component-Configurations) +- [Storage tables](https://api.keboola.com/?service=storage#tag--Tables) +- [File uploads](https://api.keboola.com/?service=storage#tag--Files) +- [Storage buckets](https://api.keboola.com/?service=storage#tag--Buckets) + +Virtually, all API calls require a [Storage API token](/storage/tokens/) to +be passed as the `X-StorageApi-Token` header. +Please note that the Storage API calls require the request to be sent +as `form-data` (unlike the rest of Keboola API, which is sent as `application/json`). + +For exporting tables from and importing tables to Storage, we highly recommend that you use one of the +[available clients](/storage/api/clients/) or the [Storage API Importer service](/storage/api/importer/). +All imports and exports are done using CSV files. See +the [RFC4180 Specification](https://tools.ietf.org/html/rfc4180) for the format +and encoding specification, and +[User documentation](/storage/tables/csv-files/) for help on how to create such files. + +Continue reading the following sections for guidance on how to get started: + +- [Storage importer service for the easiest upload of data via API](/storage/api/importer/) +- [Getting started with component configurations](/storage/api/configurations/) +- [Importing and exporting data](/storage/api/import-export/) +- [TDE exporter for exporting data to Tableau Data Extracts](/storage/api/tde-exporter/) diff --git a/src/content/docs/storage/api/tde-exporter/index.md b/src/content/docs/storage/api/tde-exporter/index.md new file mode 100644 index 000000000..dd7742dec --- /dev/null +++ b/src/content/docs/storage/api/tde-exporter/index.md @@ -0,0 +1,99 @@ +--- +title: TDE Exporter +slug: 'storage/api/tde-exporter' +redirect_from: + - /integrate/storage/api/tde-exporter/ +--- + + +:::note[Legacy format] +TDE (Tableau Data Extract) is a legacy Tableau format, superseded by Hyper (`.hyper`) in newer Tableau versions. For current Tableau output, prefer the maintained Tableau [writers](/components/writers/). +::: + +[TDE Exporter](https://github.com/keboola/tde-exporter) exports tables from Keboola Storage into the +[TDE file format (Tableau Data Extract)](https://www.tableau.com/about/blog/2014/7/understanding-tableau-data-extracts-part1). +This component is normally a part of the [Tableau Writer](/tutorial/write/), +but it can also be used as a standalone component. + +Users can [run a TDE exporter job](/integrate/jobs/) as any other Keboola component or register it +as an orchestration task. After the exporter finishes, the resulting TDE files will be available in the +*Storage* --- *File uploads* section where you can download them via UI or [API](/storage/api/import-export/). + +## Running the Component +The TDE Exporter is a Keboola [component](/extend/component/) supporting both +[stored](/storage/api/configurations/) and +custom configurations supplied directly in the `run` request. + +### Stored Configuration +To run the TDE exporter with a stored configuration, first +[create the configuration](https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/components/-componentId-/configs). +See [below](#custom-configuration) for the required configuration contents. +This call will give you the ID of the newly created configuration (for instance, `new-configuration-id`). +Then [create a job](/integrate/jobs/) with the specified configuration: + +```json +{ + "config": "new-configuration-id" +} +``` + +### Custom Configuration +You can specify the entire configuration in the API call. The JSON configuration conforms +to the [general configuration format](/extend/common-interface/config-file/). The specific part +is only the `parameters` section. A sample request to the `in.c-main.old-table` export table would look like this: + +```json +{ + "configData": { + "storage": { + "input": { + "tables": [{ + "source": "in.c-main.old-table" + }] + } + }, + "parameters": { + "tags": ["sometag"], + "typedefs": { + "in.c-main.old-table": { + "id": { + "type": "number" + }, + "col1": { + "type": "string" + } + } + } + } + } +} +``` + +The `parameters` section contains: + +- `tags`: array of tags that will be assigned to the resulting file in Storage File Uploads. +- `typedefs`: definitions of data types mapping source tables columns to destination TDE columns. + +The type definitions are entered as an object whose name must match the name of the table in the +`storage.input.tables.source` node (`in.c-main.old-table` in the above example). Object properties +are names of the table columns; each must have the `type` property which is one of the +[supported column types](https://help.tableau.com/current/pro/desktop/en-us/datafields_typesandroles_datatypes.htm): +`boolean`, `number`, `decimal`, `date`, `datetime` and `string`. + +## Date and DateTime +Data for these data types can be specified in the format used +in the [strptime function](https://pubs.opengroup.org/onlinepubs/009695399/functions/strptime.html). The format is specified as part of the column's type definition. For example: + +```json +{ + "col1": { + "type": "date", + "format":"%m-%d-%Y" + } +} +``` + +If no format is specified, the following default formats are used: + +- For `date`: `%Y-%m-%d` +- For `datetime`: `%Y-%m-%d %H:%M:%S or %Y-%m-%d %H:%M:%S.%f` diff --git a/src/sidebar.mjs b/src/sidebar.mjs index c2465d406..4f82e7f61 100644 --- a/src/sidebar.mjs +++ b/src/sidebar.mjs @@ -444,6 +444,21 @@ export const sidebar = [ ], }, { slug: "storage/byobq" }, + { + label: "Storage API", + collapsed: true, + items: [ + { label: "Overview", slug: "storage/api" }, + { slug: "storage/api/configurations" }, + { slug: "storage/api/import-export" }, + { slug: "storage/api/importer" }, + { slug: "storage/api/tde-exporter" }, + { slug: "storage/api/clients/python-client" }, + { slug: "storage/api/clients/r-client" }, + { slug: "storage/api/clients/php-client" }, + { slug: "storage/api/clients/docker-cli" }, + ], + }, ], }, {