Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 32 additions & 0 deletions _data/navigation.yml
Original file line number Diff line number Diff line change
Expand Up @@ -736,3 +736,35 @@ items:
title: Run Orchestration
- url: /automate/set-schedule/
title: Set Schedule
- url: /integrate/
title: Integration
items:
- url: /integrate/storage/api/
title: Storage API
items:
- url: /integrate/storage/api/configurations/
title: Configurations
- url: /integrate/storage/api/import-export/
title: Import & Export
- url: /integrate/storage/api/importer/
title: API Importer
- url: /integrate/storage/api/tde-exporter/
title: TDE Exporter
- url: /integrate/storage/python-client/
title: Storage API Python Client
- url: /integrate/storage/r-client/
title: Storage API R Client
- url: /integrate/storage/php-client/
title: Storage API PHP Client
- url: /integrate/storage/docker-cli-client/
title: Storage API Docker CLI Client
- url: /integrate/variables/
title: Variables
items:
- url: /integrate/variables/tutorial/
title: Variables Tutorial
- url: /integrate/artifacts/
title: Artifacts
items:
- url: /integrate/artifacts/tutorial/
title: Artifacts Tutorial
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
1 change: 1 addition & 0 deletions public/integrate/storage/api/async-import-handling.svg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
5 changes: 5 additions & 0 deletions public/integrate/storage/new-table.csv
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
"id","secondCol"
"1","a"
"2","b"
"3","c"
"4","d"
21 changes: 21 additions & 0 deletions public/integrate/variables/countries.csv
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
"COUNTRY","CARS"
"Belgium","6293781"
"Finland","3358232"
"Italy","41393877"
"Romania","6541260"
"Turkey","20193915"
"Bulgaria","2823705"
"France","38720798"
"Netherlands","8977994"
"Russia","42201083"
"Ukraine","8655700"
"Czech Republic","5116750"
"Germany","47418800"
"Poland","20671278"
"Spain","27528877"
"United Kingdom","33792233"
"Azerbaijan","1080912"
"Denmark","2723040"
"Hungary","3393075"
"Portugal","5650428"
"Sweden","5126572"
Binary file added public/integrate/variables/tutorial-1.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added public/integrate/variables/tutorial-2.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
3 changes: 3 additions & 0 deletions public/integrate/variables/variables.svg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
1 change: 1 addition & 0 deletions src/content/docs/ai/mcp-server/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@ title: Keboola Model Context Protocol (MCP) Server
slug: 'ai/mcp-server'
redirect_from:
- /external-integrations/mcp-server/
- /integrate/mcp/
---

:::caution
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,6 @@ If you check the **Wait for result** option, the component will wait for the job
You can find out how to get a service account token in the [dbt Cloud documentation](https://docs.getdbt.com/docs/dbt-cloud-apis/service-tokens).

## Notes on Artifacts Usage
In order to be able to use Keboola artifacts, the project must have the ```artifact``` feature enabled. You can find more information about this in [Keboola's docs](https://developers.keboola.com/integrate/artifacts/).
In order to be able to use Keboola artifacts, the project must have the ```artifact``` feature enabled. You can find more information about this in [Keboola's docs](/integrate/artifacts/).


1 change: 1 addition & 0 deletions src/content/docs/components/extractors/database/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@ title: Database Data Source Connectors
slug: 'components/extractors/database'
redirect_from:
- /extractors/database/
- /integrate/database/

---

Expand Down
2 changes: 1 addition & 1 deletion src/content/docs/components/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -133,7 +133,7 @@ The bottom right panel shows a list of the configuration versions. Use the list
- roll back to an older version.

All of the operations can be [accessed via an API](https://api.keboola.com/?service=storage#post-/v2/storage/branch/-branchId-/components/-componentId-/configs).
The [developer guide](https://developers.keboola.com/integrate/storage/api/configurations/) explains how to work with configurations.
The [developer guide](/integrate/storage/api/configurations/) explains how to work with configurations.

**Important**: Component configurations do not count towards your project quota.

Expand Down
2 changes: 1 addition & 1 deletion src/content/docs/data-apps/python-js/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -327,7 +327,7 @@ def load_table(table_id: str) -> pd.DataFrame:
return pd.read_csv(StringIO(response.text))
```

For a complete example using the official Python client library, see the [Keboola Storage Python Client documentation](https://developers.keboola.com/integrate/storage/python-client/).
For a complete example using the official Python client library, see the [Keboola Storage Python Client documentation](/integrate/storage/python-client/).

## Secrets and Environment Variables

Expand Down
1 change: 1 addition & 0 deletions src/content/docs/flows/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@ title: Conditional Flows
slug: 'flows'
redirect_from:
- /flows/conditional-flows/
- /integrate/orchestrator/
---

Flows allow you to build automated data pipelines with conditional logic, branching, retries, and robust error handling. You can define flows that react to the outcome of previous steps, dynamically control their next action, or even skip tasks entirely.
Expand Down
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
129 changes: 129 additions & 0 deletions src/content/docs/integrate/artifacts/index.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,129 @@
---
title: Artifacts
slug: 'integrate/artifacts'
---


:::caution[Public Beta]
This is a preview feature and may change considerably in the future. The project must have the `artifacts` feature enabled.
:::

**Artifacts** are additional files that can be produced or consumed by a [component](https://developers.keboola.com/extend/component).

See the [Tutorial](/integrate/artifacts/tutorial/) for a step-by-step example.

## Introduction
In some cases it's useful if a component not only extracts, transforms or uploads data, but also generate some other output, metadata or other runtime-discovered data.
These could be for example:
- AI models
- performance graphs of such models
- status updates from long-running tasks
- documentation
- data quality checks from in-progress tasks

These additional information can be stored in artifacts and processed by another component or 3rd party tool.

## Storage
Artifacts are stored in Keboola File Storage.

## Types of artifacts
There are three types of artifacts for now `runs`, `custom` and `shared`.
The type specifies which components will have access to the artifact or which artifacts to download for the component to process.
Types are used in a configuration of a consumer component to specify which artifacts to download.

- **runs** - artifacts from previous runs of the same configuration

- **custom** - artifacts from previous runs of a different configuration. The configuration which produced the artifacts will be defined in the consumer configuration (configurationId, componentId, branchId)

- **shared** - artifacts shared within an orchestration

`runs` and `custom` types are the same from the producer point of view. To produce a `shared` artifact, it has to be written into a `shared` folder. Read more in [File structure](#file-structure) section.

## File structure
Artifact is a unique set of files associated with a successful job, component and configuration.
A component can either produce or consume artifacts or both.

### Produce
To produce an artifact, store one or more files in the following `output` directories. Subdirectories are also supported.
- `/data/artifacts/out/current` to create an artifact of type `runs` / `custom`.
- `/data/artifacts/out/shared` to create an artifact of type `shared`, which can be accessed by any component within the same orchestration.

After the component job is finished all files and directories inside `current` and `shared` folders will be compressed into an archive and uploaded to File Storage with corresponding tags as a `artifact`.

### Consume
To consume created artifacts you have to specify, in the configuration of a component, which artifacts (type) to download.
- `runs` to download artifacts produced by the same configuration and component. These will be stored in `/data/artifacts/in/runs/jobs/job-%job_id%` directory.
- `custom` to download artifacts produced by another configuration or component. These will be stored in `/data/artifacts/in/custom/jobs/job-%job_id%` directory.
- `shared` to download artifacts created within the same orchestration by any artifact producing component that has already finished. These will be stored in `/data/artifacts/in/shared/jobs/job-%job_id%` directory.

## Configuration
Each type of artifact has a separate node in configuration. All the types can be used simultaneously.
Each type node has an attribute "enabled", which enables or disables download of the corresponding artifact type.

### Runs
- **enabled** [true|false] - enable or disable download of this artifact type
- **filter**
- **date_since** - only artifacts from jobs younger than this will be downloaded
- **limit** - maximum number of the latest jobs from which to download artifacts

### Custom
- **enabled** [true|false] - enable or disable download of this artifact type
- **filter**
- **branch_id**, **component_id**, **config_id** - specify the configuration to download artifacts from
- **date_since** - only artifacts from jobs younger than this will be downloaded
- **limit** - maximum number of the latest jobs from which to download artifacts

### Shared
- **enabled** [true|false] - enable or disable download of this artifact type

Full configuration example with all artifact types:

```json
{
"parameters": {},
"artifacts": {
"runs": {
"enabled": true,
"filter": {
"date_since": "-7 days",
"limit": 5
}
},
"custom": {
"enabled": true,
"filter": {
"component_id": "keboola.python-transformation",
"config_id": "12345",
"branch_id": "default",
"date_since": "-7 days",
"limit": 5
}
},
"orchestration": {
"enabled": true
}
}
}
```

## Artifacts life-cycle in a job
Job runner checks if the project has enabled `artifacts` feature.
Job runner checks the configuration of the component.
If artifacts are enabled, it downloads artifacts to corresponding folders as configured (i.e. `runs`, `custom`, `shared`) and unzips them.

Component process start and the component can:

- access and process the downloaded artifacts in shared or custom directory

- write artifacts to `current` or `shared` directory

Component finishes and job runner does:

- gzip the content of runs/current

- tag the gzipped file with jobId, componentId, configId, runId, branchId and other tags if needed

- upload the file to File Storage

## File size limit
All the artifacts produced by a job shouldn’t be bigger than 1 GB.
Loading