From 861f13ed925ba5902c6f309fb82f47073e29a6f6 Mon Sep 17 00:00:00 2001 From: Nikita Date: Thu, 9 Jul 2026 15:09:02 +0200 Subject: [PATCH 1/2] docs(tutorial): flag legacy-UI walkthroughs; fix Conditional Flows label MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit automate: the walkthrough screenshots show the Legacy Flow Builder — added a caution pointing to Conditional Flows (nav label verified live) and a TODO(human-review) to re-record on the conditional builder. ad-hoc: flagged the 'Exploring Data' part for rewrite — it still shows the retired 'New Sandbox' flow (page's own caution admits this). Both pages also get description frontmatter. Co-Authored-By: Claude Opus 4.8 --- src/content/docs/tutorial/ad-hoc/index.md | 510 ++++++++++---------- src/content/docs/tutorial/automate/index.md | 8 +- 2 files changed, 263 insertions(+), 255 deletions(-) diff --git a/src/content/docs/tutorial/ad-hoc/index.md b/src/content/docs/tutorial/ad-hoc/index.md index f48296401..9132c8d2c 100644 --- a/src/content/docs/tutorial/ad-hoc/index.md +++ b/src/content/docs/tutorial/ad-hoc/index.md @@ -1,264 +1,266 @@ --- title: "Part 5: Ad-Hoc Data Analysis" slug: 'tutorial/ad-hoc' ---- - -After you have loaded your tables, either [manually](/tutorial/load/) or -[using a data source connector](/tutorial/load/database/), [manipulated the data](/tutorial/manipulate/) in SQL, -written it [into Google Sheets](/tutorial/write/), and -set everything to run [automatically](/tutorial/automate/), let's take a look at some additional Keboola -features related to doing ad-hoc analysis. - -This part of the tutorial shows how to work with arbitrary data in Python -in a completely unrestricted way. Although our examples use the Python language, -the very same can be achieved using R. - -Before you start, you should have a basic understanding of the [Python language](https://www.python.org/). - - - -## Introduction -Let's say you want to experiment with the US unemployment data. It is provided by the -[U.S. Bureau of Labor Statistics](https://www.bls.gov/cps/tables.htm) (BLS), and the dataset [A-10](https://www.bls.gov/web/empsit/cpseea10.htm) -contains unemployment rates by month. The easiest way to access the data is via -[Google Public Data](https://cloud.google.com/bigquery/public-data/), which contains a dataset called -[Bureau of Labor Statistics Data](https://cloud.google.com/bigquery/public-data/bureau-of-labor-statistics). - -Google Public Data can be queried using [BigQuery](https://cloud.google.com/bigquery/) and brought into Keboola -with the help of our BigQuery data source connector. Preview the table data in -[Google BigQuery](https://bigquery.cloud.google.com/table/bigquery-public-data:bls.unemployment_cps?tab=preview). - -## Using BigQuery Connector -To work with Google BigQuery, create an account, and [enable billing](https://cloud.google.com/bigquery/public-data/). Remember, -querying public data is only [free up to 1TB a month](https://cloud.google.com/bigquery/public-data/). - -Then create a [service account](https://cloud.google.com/bigquery/docs/authentication/#service_accounts) for authentication -of the Google BigQuery data source connector, and create a Google Storage bucket as a temporary storage for off-loading the data from BigQuery. - -***Note:** If setting up the Google BigQuery connector seems too complicated to you, export the query results to Google Sheets and -[load them from Google Sheets](/tutorial/load/googlesheets/). Or, export them to a CSV file and [load them from local files](/tutorial/load/#manually-loading-data).* - -### Prepare -Before you start, have a Google service account and a Google Storage bucket ready. - -#### Service account -To create a Google service account, go to the -[**Google Cloud Platform Console > IAM & admin > Service accounts**](https://console.cloud.google.com/iam-admin/serviceaccounts) -and create a new service account: - -![Screenshot - Google Service Account](/tutorial/ad-hoc/cloud-platform-service-account-1.png) - -Name the service account: - -![Screenshot - Google Service Account Detail](/tutorial/ad-hoc/cloud-platform-service-account-3.png) - -Grant the roles **BigQuery Data Editor**, **BigQuery Job User** and **Storage Object Admin** to your service account: - -![Screenshot - Google Service Account Permissions](/tutorial/ad-hoc/cloud-platform-service-account-4.png) - -Finally, create a new JSON key and download it to your computer: - -![Screenshot - Google Service Account Download](/tutorial/ad-hoc/cloud-platform-service-account-5.png) - -#### Google Storage bucket -To create a Google Storage bucket, go to the [**Google Cloud Platform console > Storage**](https://console.cloud.google.com/storage/browser) -and create a new bucket: - -![Screenshot - Google Cloud Platform](/tutorial/ad-hoc/cloud-platform-storage-1.png) - -Enter the bucket's name and choose where to store your data (the location type *Region* is okay for our purpose): - -![Screenshot - Create Bucket](/tutorial/ad-hoc/cloud-platform-storage-3.png) - -Do not set a retention policy on the bucket. The bucket contains only temporary data and no retention is needed. - -### Extract Data -Now you're ready to load the data into Keboola. Go to the section **Components**, -and click the green button **Add Component**: - -![Screenshot - Extractors](/tutorial/ad-hoc/ex-bigquery-1.png) - -Use the search to find the Google BigQuery data source: - -![Screenshot - BigQuery Extractor](/tutorial/ad-hoc/ex-bigquery-2.png) - -Click **+ Add Component** and then **Connect to My Data**: - -![Screenshot - New Configuration](/tutorial/ad-hoc/ex-bigquery-3.png) - -Name the configuration (e.g., 'Bls Unemployment') and describe it if you want. Then, click **Create Configuration**: - -![Screenshot - New Configuration Name](/tutorial/ad-hoc/ex-bigquery-4.png) - -Then set the service account key: - -![Screenshot - Big Query Authorization](/tutorial/ad-hoc/ex-bigquery-5.png) - -Open the downloaded key you have created above in a text editor, copy & paste it in the input field, click **Submit** and then **Save**. - -![Screenshot - Service Account Copy](/tutorial/ad-hoc/ex-bigquery-6.png) - -Fill the bucket you have created above: - -![Screenshot - Big Query Unload](/tutorial/ad-hoc/ex-bigquery-7.png) - -After that configure the actual extraction queries by clicking the **Add Query** button: - -![Screenshot - Big Query Configured](/tutorial/ad-hoc/ex-bigquery-8.png) - -Name the query, e.g., `Unemployment rates`: - -![Screenshot - New Query Name](/tutorial/ad-hoc/ex-bigquery-9.png) - -Check *Create your own query using an SQL editor*, uncheck the *Use Legacy SQL* setting, and paste the following code in the *SQL Query* field: - - +description: 'Part 5 of the tutorial — pull public data with the BigQuery connector and explore it ad hoc with Python, pandas, and Matplotlib in a Jupyter workspace.' +--- + +After you have loaded your tables, either [manually](/tutorial/load/) or +[using a data source connector](/tutorial/load/database/), [manipulated the data](/tutorial/manipulate/) in SQL, +written it [into Google Sheets](/tutorial/write/), and +set everything to run [automatically](/tutorial/automate/), let's take a look at some additional Keboola +features related to doing ad-hoc analysis. + +This part of the tutorial shows how to work with arbitrary data in Python +in a completely unrestricted way. Although our examples use the Python language, +the very same can be achieved using R. + +Before you start, you should have a basic understanding of the [Python language](https://www.python.org/). + + + +## Introduction +Let's say you want to experiment with the US unemployment data. It is provided by the +[U.S. Bureau of Labor Statistics](https://www.bls.gov/cps/tables.htm) (BLS), and the dataset [A-10](https://www.bls.gov/web/empsit/cpseea10.htm) +contains unemployment rates by month. The easiest way to access the data is via +[Google Public Data](https://cloud.google.com/bigquery/public-data/), which contains a dataset called +[Bureau of Labor Statistics Data](https://cloud.google.com/bigquery/public-data/bureau-of-labor-statistics). + +Google Public Data can be queried using [BigQuery](https://cloud.google.com/bigquery/) and brought into Keboola +with the help of our BigQuery data source connector. Preview the table data in +[Google BigQuery](https://bigquery.cloud.google.com/table/bigquery-public-data:bls.unemployment_cps?tab=preview). + +## Using BigQuery Connector +To work with Google BigQuery, create an account, and [enable billing](https://cloud.google.com/bigquery/public-data/). Remember, +querying public data is only [free up to 1TB a month](https://cloud.google.com/bigquery/public-data/). + +Then create a [service account](https://cloud.google.com/bigquery/docs/authentication/#service_accounts) for authentication +of the Google BigQuery data source connector, and create a Google Storage bucket as a temporary storage for off-loading the data from BigQuery. + +***Note:** If setting up the Google BigQuery connector seems too complicated to you, export the query results to Google Sheets and +[load them from Google Sheets](/tutorial/load/googlesheets/). Or, export them to a CSV file and [load them from local files](/tutorial/load/#manually-loading-data).* + +### Prepare +Before you start, have a Google service account and a Google Storage bucket ready. + +#### Service account +To create a Google service account, go to the +[**Google Cloud Platform Console > IAM & admin > Service accounts**](https://console.cloud.google.com/iam-admin/serviceaccounts) +and create a new service account: + +![Screenshot - Google Service Account](/tutorial/ad-hoc/cloud-platform-service-account-1.png) + +Name the service account: + +![Screenshot - Google Service Account Detail](/tutorial/ad-hoc/cloud-platform-service-account-3.png) + +Grant the roles **BigQuery Data Editor**, **BigQuery Job User** and **Storage Object Admin** to your service account: + +![Screenshot - Google Service Account Permissions](/tutorial/ad-hoc/cloud-platform-service-account-4.png) + +Finally, create a new JSON key and download it to your computer: + +![Screenshot - Google Service Account Download](/tutorial/ad-hoc/cloud-platform-service-account-5.png) + +#### Google Storage bucket +To create a Google Storage bucket, go to the [**Google Cloud Platform console > Storage**](https://console.cloud.google.com/storage/browser) +and create a new bucket: + +![Screenshot - Google Cloud Platform](/tutorial/ad-hoc/cloud-platform-storage-1.png) + +Enter the bucket's name and choose where to store your data (the location type *Region* is okay for our purpose): + +![Screenshot - Create Bucket](/tutorial/ad-hoc/cloud-platform-storage-3.png) + +Do not set a retention policy on the bucket. The bucket contains only temporary data and no retention is needed. + +### Extract Data +Now you're ready to load the data into Keboola. Go to the section **Components**, +and click the green button **Add Component**: + +![Screenshot - Extractors](/tutorial/ad-hoc/ex-bigquery-1.png) + +Use the search to find the Google BigQuery data source: + +![Screenshot - BigQuery Extractor](/tutorial/ad-hoc/ex-bigquery-2.png) + +Click **+ Add Component** and then **Connect to My Data**: + +![Screenshot - New Configuration](/tutorial/ad-hoc/ex-bigquery-3.png) + +Name the configuration (e.g., 'Bls Unemployment') and describe it if you want. Then, click **Create Configuration**: + +![Screenshot - New Configuration Name](/tutorial/ad-hoc/ex-bigquery-4.png) + +Then set the service account key: + +![Screenshot - Big Query Authorization](/tutorial/ad-hoc/ex-bigquery-5.png) + +Open the downloaded key you have created above in a text editor, copy & paste it in the input field, click **Submit** and then **Save**. + +![Screenshot - Service Account Copy](/tutorial/ad-hoc/ex-bigquery-6.png) + +Fill the bucket you have created above: + +![Screenshot - Big Query Unload](/tutorial/ad-hoc/ex-bigquery-7.png) + +After that configure the actual extraction queries by clicking the **Add Query** button: + +![Screenshot - Big Query Configured](/tutorial/ad-hoc/ex-bigquery-8.png) + +Name the query, e.g., `Unemployment rates`: + +![Screenshot - New Query Name](/tutorial/ad-hoc/ex-bigquery-9.png) + +Check *Create your own query using an SQL editor*, uncheck the *Use Legacy SQL* setting, and paste the following code in the *SQL Query* field: + + ```sql -SELECT * FROM - `bigquery-public-data.bls.unemployment_cps` -WHERE - series_id = "LNS14000000" -ORDER BY date -``` - -The `LNS14000000` series will pick the unemployment rates only. - -Then **Save** the query configuration. - -![Screenshot - Query Configuration](/tutorial/ad-hoc/ex-bigquery-10.png) - -Now run the configuration to bring the data to Keboola: - -![Screenshot - Finished Configuration](/tutorial/ad-hoc/ex-bigquery-11.png) - -Running the data source connector creates a background job that - -- executes the queries in Google BigQuery. -- saves the results to Google Cloud Storage. -- exports the results from Google Cloud Storage and stores them in specified tables in Keboola Storage. -- removes the results from Google Cloud Storage. - -When a job is running, a small orange circle appears under *Last runs*, along with RunId and other info on the job. -Green is for success, red for failure. Click on the indicator, or the info next to it, for more details. -Once the job is finished, click on the names of the tables to inspect their contents. - -## Exploring Data - +SELECT * FROM + `bigquery-public-data.bls.unemployment_cps` +WHERE + series_id = "LNS14000000" +ORDER BY date +``` + +The `LNS14000000` series will pick the unemployment rates only. + +Then **Save** the query configuration. + +![Screenshot - Query Configuration](/tutorial/ad-hoc/ex-bigquery-10.png) + +Now run the configuration to bring the data to Keboola: + +![Screenshot - Finished Configuration](/tutorial/ad-hoc/ex-bigquery-11.png) + +Running the data source connector creates a background job that + +- executes the queries in Google BigQuery. +- saves the results to Google Cloud Storage. +- exports the results from Google Cloud Storage and stores them in specified tables in Keboola Storage. +- removes the results from Google Cloud Storage. + +When a job is running, a small orange circle appears under *Last runs*, along with RunId and other info on the job. +Green is for success, red for failure. Click on the indicator, or the info next to it, for more details. +Once the job is finished, click on the names of the tables to inspect their contents. + +## Exploring Data + :::caution **Important:** The following part of the tutorial will be updated soon. Please be aware that sandboxes now exist only in their legacy form and have been replaced by workspaces. -::: - -To explore the data, go to [**Workspaces**](/workspace/). -Provided for each user and project automatically, it is an isolated environment in which you can experiment without -interfering with any production code. - -![Screenshot - Transformations](/tutorial/ad-hoc/transformation-1.png) - -Click on **New Sandbox** next to Python (Jupyter): - -![Screenshot - Create Sandbox](/tutorial/ad-hoc/transformation-2.png) - -Select the unemployment rates table (`in.c-keboola-ex-google-bigquery-v2-548939034.unemployment-rates` in this case), -click on **Create Sandbox**. Wait for the process to finish: - -![Screenshot - Sandbox Configuration](/tutorial/ad-hoc/transformation-3.png) - -When finished, connect to the web version of the [Jupyter Notebook](http://jupyter.org/). -It allows you to run arbitrary code by clicking the **Connect** button: - -![Screenshot - Sandbox Credentials](/tutorial/ad-hoc/transformation-4.png) - -When prompted, enter the password from the Sandbox screen: - -![Screenshot - Sandbox Login](/tutorial/ad-hoc/sandbox-1.png) - -You can now run arbitrary code in Python, using common data scientist tools like -[Pandas](https://pandas.pydata.org/) or [Matplotlib](https://matplotlib.org/). -For instance, to load the file, use (make sure to use the correct filename): - +::: + + +To explore the data, go to [**Workspaces**](/workspace/). +Provided for each user and project automatically, it is an isolated environment in which you can experiment without +interfering with any production code. + +![Screenshot - Transformations](/tutorial/ad-hoc/transformation-1.png) + +Click on **New Sandbox** next to Python (Jupyter): + +![Screenshot - Create Sandbox](/tutorial/ad-hoc/transformation-2.png) + +Select the unemployment rates table (`in.c-keboola-ex-google-bigquery-v2-548939034.unemployment-rates` in this case), +click on **Create Sandbox**. Wait for the process to finish: + +![Screenshot - Sandbox Configuration](/tutorial/ad-hoc/transformation-3.png) + +When finished, connect to the web version of the [Jupyter Notebook](http://jupyter.org/). +It allows you to run arbitrary code by clicking the **Connect** button: + +![Screenshot - Sandbox Credentials](/tutorial/ad-hoc/transformation-4.png) + +When prompted, enter the password from the Sandbox screen: + +![Screenshot - Sandbox Login](/tutorial/ad-hoc/sandbox-1.png) + +You can now run arbitrary code in Python, using common data scientist tools like +[Pandas](https://pandas.pydata.org/) or [Matplotlib](https://matplotlib.org/). +For instance, to load the file, use (make sure to use the correct filename): + ```python -import pandas -df = pandas.read_csv("/data/in/tables/in.c-keboola-ex-google-bigquery-v2-548939034.unemployment-rates.csv",sep=',') -df.head() -``` - -The path `/data/in/tables/` is the location for -[loaded tables](/transformations/python-plain/#file-locations); they -are loaded as simple CSV files. Once your table is loaded, you can play with it: - +import pandas +df = pandas.read_csv("/data/in/tables/in.c-keboola-ex-google-bigquery-v2-548939034.unemployment-rates.csv",sep=',') +df.head() +``` + +The path `/data/in/tables/` is the location for +[loaded tables](/transformations/python-plain/#file-locations); they +are loaded as simple CSV files. Once your table is loaded, you can play with it: + ```python -import matplotlib.pyplot as plt -years = df.groupby(df['year'])['value'].mean() -years.plot(kind='line', color = 'orange') -plt.xlabel("Year") -plt.ylabel("Average %") -plt.suptitle('US Unemployment Rate', size=15) -plt.show() -``` - -![Screenshot - Sandbox Result](/tutorial/ad-hoc/sandbox-2.png) - -## Adding Libraries -Now that you can experiment with the U.S. unemployment data extracted from Google BigQuery (or any other data extracted in any other way), -you can do the same with the EU unemployment data. Available at [Eurostat](https://ec.europa.eu/eurostat/), the unemployment -dataset is called -[`tgs00010`](https://ec.europa.eu/eurostat/databrowser/product/view/tgs00010?lang=en). - -There are a number of ways how to get the data from Eurostat -- e.g., you can download it in TSV -or XLS format. To avoid downloading the (possibly) lengthy data set to your hard drive, Eurostat provides a -[REST API](https://ec.europa.eu/eurostat/web/user-guides/data-browser/api-data-access/api-migrating/json) -for downloading the data. This could be processed using the -[Generic Extractor](/components/extractors/other/generic/). However, the data is provided in -[JSON-stat](https://json-stat.org/) format, which contains tables encoded using the -[row-major](https://en.wikipedia.org/wiki/Row-_and_column-major_order) method. Even though it is possible -to import them to Keboola, it would be necessary to do additional processing to obtain plain tables. - -To save time, use a tool designed for that -- [pyjstat](https://pypi.org/project/pyjstat/). It is a Python library which can read -JSON-stat data directly into a [Pandas data frame](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.html). -Although this library is not installed by default in the Jupyter Sandbox environment, nothing prevents you from installing it. - -### Working with Custom Libraries -Use the following code to download the desired data from Eurostat: - +import matplotlib.pyplot as plt +years = df.groupby(df['year'])['value'].mean() +years.plot(kind='line', color = 'orange') +plt.xlabel("Year") +plt.ylabel("Average %") +plt.suptitle('US Unemployment Rate', size=15) +plt.show() +``` + +![Screenshot - Sandbox Result](/tutorial/ad-hoc/sandbox-2.png) + +## Adding Libraries +Now that you can experiment with the U.S. unemployment data extracted from Google BigQuery (or any other data extracted in any other way), +you can do the same with the EU unemployment data. Available at [Eurostat](https://ec.europa.eu/eurostat/), the unemployment +dataset is called +[`tgs00010`](https://ec.europa.eu/eurostat/databrowser/product/view/tgs00010?lang=en). + +There are a number of ways how to get the data from Eurostat -- e.g., you can download it in TSV +or XLS format. To avoid downloading the (possibly) lengthy data set to your hard drive, Eurostat provides a +[REST API](https://ec.europa.eu/eurostat/web/user-guides/data-browser/api-data-access/api-migrating/json) +for downloading the data. This could be processed using the +[Generic Extractor](/components/extractors/other/generic/). However, the data is provided in +[JSON-stat](https://json-stat.org/) format, which contains tables encoded using the +[row-major](https://en.wikipedia.org/wiki/Row-_and_column-major_order) method. Even though it is possible +to import them to Keboola, it would be necessary to do additional processing to obtain plain tables. + +To save time, use a tool designed for that -- [pyjstat](https://pypi.org/project/pyjstat/). It is a Python library which can read +JSON-stat data directly into a [Pandas data frame](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.html). +Although this library is not installed by default in the Jupyter Sandbox environment, nothing prevents you from installing it. + +### Working with Custom Libraries +Use the following code to download the desired data from Eurostat: + ```python -import subprocess -import sys -subprocess.call([sys.executable, '-m', 'pip', 'install', '--disable-pip-version-check', '-q', 'pyjstat']) -from pyjstat import pyjstat -dataset = pyjstat.Dataset.read('https://ec.europa.eu/eurostat/api/dissemination/statistics/1.0/data/tgs00010?format=JSON&unit=PC&isced11=ED0-2&isced11=ED3_4&isced11=ED5-8&isced11=NRP&isced11=TOTAL&isced11=UNK&sex=F&sex=M&sex=T&age=Y15-74&lang=EN') -df = dataset.write('dataframe') -df.head() -``` - -The URL was built using the Eurostat [Query Builder](https://ec.europa.eu/eurostat/web/query-builder/tool). -Also note that installing a library from within the Python code must be done using `pip install`. Now that you have the data, -feel free to play with it: - +import subprocess +import sys +subprocess.call([sys.executable, '-m', 'pip', 'install', '--disable-pip-version-check', '-q', 'pyjstat']) +from pyjstat import pyjstat +dataset = pyjstat.Dataset.read('https://ec.europa.eu/eurostat/api/dissemination/statistics/1.0/data/tgs00010?format=JSON&unit=PC&isced11=ED0-2&isced11=ED3_4&isced11=ED5-8&isced11=NRP&isced11=TOTAL&isced11=UNK&sex=F&sex=M&sex=T&age=Y15-74&lang=EN') +df = dataset.write('dataframe') +df.head() +``` + +The URL was built using the Eurostat [Query Builder](https://ec.europa.eu/eurostat/web/query-builder/tool). +Also note that installing a library from within the Python code must be done using `pip install`. Now that you have the data, +feel free to play with it: + ```python -years = df.groupby(df['time'])['value'].mean() -years.plot(kind='line', color = 'orange') -plt.xlabel("Year") -plt.ylabel("Average %") -plt.suptitle('EU Unemployment Rate', size=15) -plt.show() -``` - -![Screenshot - Sandbox Result](/tutorial/ad-hoc/sandbox-3.png) - -## Wrap Up -You have just learnt to do a completely ad-hoc analysis of various data sets. If you need to run the above code regularly, -simply copy&paste it into a [Transformation](/tutorial/manipulate/). - -The above tutorial is done in the [Python language](https://www.python.org/) using the -[Jupyter Notebook](https://jupyter.org/). The same can be done in the -[R language](https://www.r-project.org/) using [RStudio](https://rstudio.com/). -For more information about workspaces (including disk and memory limits), see the -[corresponding documentation](/workspace/). - -## Final Note -This is the end of our stroll around Keboola. On our walk, we missed quite a few things: -Applications, Python and R transformations, Snowflake features, to name a few. -However, teaching you everything was not really the point of this tutorial. -We wanted to show you how Keboola can help in connecting different systems together. - -[Return to the beginning](/tutorial/) or [contact us](/). +years = df.groupby(df['time'])['value'].mean() +years.plot(kind='line', color = 'orange') +plt.xlabel("Year") +plt.ylabel("Average %") +plt.suptitle('EU Unemployment Rate', size=15) +plt.show() +``` + +![Screenshot - Sandbox Result](/tutorial/ad-hoc/sandbox-3.png) + +## Wrap Up +You have just learnt to do a completely ad-hoc analysis of various data sets. If you need to run the above code regularly, +simply copy&paste it into a [Transformation](/tutorial/manipulate/). + +The above tutorial is done in the [Python language](https://www.python.org/) using the +[Jupyter Notebook](https://jupyter.org/). The same can be done in the +[R language](https://www.r-project.org/) using [RStudio](https://rstudio.com/). +For more information about workspaces (including disk and memory limits), see the +[corresponding documentation](/workspace/). + +## Final Note +This is the end of our stroll around Keboola. On our walk, we missed quite a few things: +Applications, Python and R transformations, Snowflake features, to name a few. +However, teaching you everything was not really the point of this tutorial. +We wanted to show you how Keboola can help in connecting different systems together. + +[Return to the beginning](/tutorial/) or [contact us](/). diff --git a/src/content/docs/tutorial/automate/index.md b/src/content/docs/tutorial/automate/index.md index 734d343be..974860978 100644 --- a/src/content/docs/tutorial/automate/index.md +++ b/src/content/docs/tutorial/automate/index.md @@ -1,6 +1,7 @@ --- title: "Part 4: Flow Automation" slug: 'tutorial/automate' +description: 'Part 4 of the tutorial — orchestrate the load, transform, and write steps into a flow and schedule it to run automatically every day.' --- So far, you have learned to use Keboola to @@ -16,7 +17,12 @@ This is where our flows come in: - Specify what tasks should be executed in what order (orchestrate tasks) and - Configure the automatic execution (schedule flow tasks). -1. Navigate to the **Flows** section of Keboola. +:::caution +The screenshots below show the **Legacy Flow Builder**. New projects default to [Conditional Flows](/flows/), which offer conditional logic, retries, and more — the concepts here (steps, parallel tasks, schedules, notifications) carry over. To build this pipeline in the Conditional Flows builder, see [Conditional Flows](/flows/). +::: + + +1. Navigate to the **Conditional Flows** section of Keboola. ![Go to Flows](/tutorial/automate/automate1.png) From 2dec6ec16486f9a63403591d02e23b537112a609 Mon Sep 17 00:00:00 2001 From: Nikita Date: Thu, 9 Jul 2026 15:09:02 +0200 Subject: [PATCH 2/2] docs(tutorial): add descriptions to all remaining tutorial pages MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 20 pages across load, manipulate, write, branches, and onboarding — every tutorial page now has description frontmatter for search/RAG/meta. Co-Authored-By: Claude Opus 4.8 --- src/content/docs/tutorial/branches/files-in-branch.md | 1 + src/content/docs/tutorial/branches/index.md | 1 + src/content/docs/tutorial/branches/merge-to-production.md | 1 + src/content/docs/tutorial/branches/prepare-files.md | 1 + src/content/docs/tutorial/branches/prepare-tables.md | 1 + src/content/docs/tutorial/branches/project-diff.md | 1 + src/content/docs/tutorial/branches/tables-in-branch.md | 1 + src/content/docs/tutorial/index.md | 1 + src/content/docs/tutorial/load/database.md | 1 + src/content/docs/tutorial/load/googlesheets.md | 1 + src/content/docs/tutorial/load/manual.md | 1 + src/content/docs/tutorial/manipulate/index.md | 1 + src/content/docs/tutorial/manipulate/workspace.md | 1 + .../tutorial/onboarding/architecture-guide/bdm-guide/index.md | 1 + src/content/docs/tutorial/onboarding/architecture-guide/index.md | 1 + src/content/docs/tutorial/onboarding/cheat-sheet/index.md | 1 + src/content/docs/tutorial/onboarding/governance-guide/index.md | 1 + src/content/docs/tutorial/onboarding/index.md | 1 + src/content/docs/tutorial/onboarding/usage-blueprint/index.md | 1 + src/content/docs/tutorial/write/index.md | 1 + 20 files changed, 20 insertions(+) diff --git a/src/content/docs/tutorial/branches/files-in-branch.md b/src/content/docs/tutorial/branches/files-in-branch.md index 9bc9ee83c..5446ef198 100644 --- a/src/content/docs/tutorial/branches/files-in-branch.md +++ b/src/content/docs/tutorial/branches/files-in-branch.md @@ -1,6 +1,7 @@ --- title: Working with Files in Branch slug: 'tutorial/branches/files-in-branch' +description: 'Learn how files behave in development branches by running the file-manipulating configurations in a branch.' --- diff --git a/src/content/docs/tutorial/branches/index.md b/src/content/docs/tutorial/branches/index.md index a1d03b0a7..a127c9308 100644 --- a/src/content/docs/tutorial/branches/index.md +++ b/src/content/docs/tutorial/branches/index.md @@ -1,6 +1,7 @@ --- title: "Part 6: Development Branches" slug: 'tutorial/branches' +description: 'Part 6 of the tutorial — use development branches to change configurations safely without touching production pipelines.' --- diff --git a/src/content/docs/tutorial/branches/merge-to-production.md b/src/content/docs/tutorial/branches/merge-to-production.md index e56a1b531..f0845f7c7 100644 --- a/src/content/docs/tutorial/branches/merge-to-production.md +++ b/src/content/docs/tutorial/branches/merge-to-production.md @@ -1,6 +1,7 @@ --- title: Merge to Production slug: 'tutorial/branches/merge-to-production' +description: 'Merge a reviewed development branch back to production and clean up after the merge.' --- diff --git a/src/content/docs/tutorial/branches/prepare-files.md b/src/content/docs/tutorial/branches/prepare-files.md index 01aaaf7ee..17ca93cf0 100644 --- a/src/content/docs/tutorial/branches/prepare-files.md +++ b/src/content/docs/tutorial/branches/prepare-files.md @@ -1,6 +1,7 @@ --- title: Prepare File Manipulating Configurations slug: 'tutorial/branches/prepare-files' +description: 'Set up the production file-manipulating configurations used in the development-branches tutorial.' --- diff --git a/src/content/docs/tutorial/branches/prepare-tables.md b/src/content/docs/tutorial/branches/prepare-tables.md index b1ecb88b7..1fc19cdc5 100644 --- a/src/content/docs/tutorial/branches/prepare-tables.md +++ b/src/content/docs/tutorial/branches/prepare-tables.md @@ -1,6 +1,7 @@ --- title: Prepare Table-Manipulating Configurations slug: 'tutorial/branches/prepare-tables' +description: 'Set up the production table-manipulating configurations used throughout the development-branches tutorial.' --- diff --git a/src/content/docs/tutorial/branches/project-diff.md b/src/content/docs/tutorial/branches/project-diff.md index bd36b43ce..5f95040f3 100644 --- a/src/content/docs/tutorial/branches/project-diff.md +++ b/src/content/docs/tutorial/branches/project-diff.md @@ -1,6 +1,7 @@ --- title: Project Diff slug: 'tutorial/branches/project-diff' +description: 'Review changes between your development branch and production with the project diff before merging.' --- diff --git a/src/content/docs/tutorial/branches/tables-in-branch.md b/src/content/docs/tutorial/branches/tables-in-branch.md index 9c768b231..9b1a3ee91 100644 --- a/src/content/docs/tutorial/branches/tables-in-branch.md +++ b/src/content/docs/tutorial/branches/tables-in-branch.md @@ -1,6 +1,7 @@ --- title: Working with Tables in Branch slug: 'tutorial/branches/tables-in-branch' +description: 'Run configurations inside a development branch and learn how Storage tables behave in branches.' --- diff --git a/src/content/docs/tutorial/index.md b/src/content/docs/tutorial/index.md index 3da71ab35..3e6d6b086 100644 --- a/src/content/docs/tutorial/index.md +++ b/src/content/docs/tutorial/index.md @@ -3,6 +3,7 @@ title: Keboola Getting Started Tutorial slug: 'tutorial' redirect_from: - /getting-started/ +description: 'A hands-on tour of Keboola — load data, transform it with SQL, write it to a destination, automate the pipeline with a flow, and explore ad-hoc analysis.' --- Discover how to leverage the Keboola platform to effortlessly extract data from various sources, transform it, and securely store it within the Keboola platform. diff --git a/src/content/docs/tutorial/load/database.md b/src/content/docs/tutorial/load/database.md index 8c449e1d3..b1150a0b0 100644 --- a/src/content/docs/tutorial/load/database.md +++ b/src/content/docs/tutorial/load/database.md @@ -1,6 +1,7 @@ --- title: Loading Data from Database slug: 'tutorial/load/database' +description: 'Load data from an external Snowflake database with a database data source connector — the procedure is the same for all database sources.' --- So far, you have learned to load data into Keboola [manually](/tutorial/load/) and diff --git a/src/content/docs/tutorial/load/googlesheets.md b/src/content/docs/tutorial/load/googlesheets.md index d96119349..b751b9e5a 100644 --- a/src/content/docs/tutorial/load/googlesheets.md +++ b/src/content/docs/tutorial/load/googlesheets.md @@ -1,6 +1,7 @@ --- title: Loading Data from Google Sheets slug: 'tutorial/load/googlesheets' +description: 'Load a shared Google spreadsheet into Storage with the Google Sheets data source connector.' --- In the [previous step](/tutorial/load/), you learned how to quickly load data into Keboola using [manual import](/tutorial/load/). diff --git a/src/content/docs/tutorial/load/manual.md b/src/content/docs/tutorial/load/manual.md index 2a82927f4..c843b3a15 100644 --- a/src/content/docs/tutorial/load/manual.md +++ b/src/content/docs/tutorial/load/manual.md @@ -1,6 +1,7 @@ --- title: "Part 1: Loading Data" slug: 'tutorial/load' +description: 'Part 1 of the tutorial — load CSV files into Storage manually, the quickest way to get data in during a proof of concept.' --- diff --git a/src/content/docs/tutorial/manipulate/index.md b/src/content/docs/tutorial/manipulate/index.md index 7ccb48d4e..42cf97a35 100644 --- a/src/content/docs/tutorial/manipulate/index.md +++ b/src/content/docs/tutorial/manipulate/index.md @@ -1,6 +1,7 @@ --- title: "Part 2: Data Manipulation" slug: 'tutorial/manipulate' +description: 'Part 2 of the tutorial — build a SQL transformation that denormalizes the loaded tables, with input and output mappings explained step by step.' --- At this juncture, you're acquainted with the swift process of loading data into Keboola, resulting in four new tables in your Storage: diff --git a/src/content/docs/tutorial/manipulate/workspace.md b/src/content/docs/tutorial/manipulate/workspace.md index 8fc26f807..66a957f85 100644 --- a/src/content/docs/tutorial/manipulate/workspace.md +++ b/src/content/docs/tutorial/manipulate/workspace.md @@ -1,6 +1,7 @@ --- title: Using a Workspace slug: 'tutorial/manipulate/workspace' +description: 'Develop and test your transformation SQL interactively in a workspace before saving it to the transformation.' --- An integral aspect of creating a transformation is the development of the script itself. diff --git a/src/content/docs/tutorial/onboarding/architecture-guide/bdm-guide/index.md b/src/content/docs/tutorial/onboarding/architecture-guide/bdm-guide/index.md index 9ae05d723..45cc54fd6 100644 --- a/src/content/docs/tutorial/onboarding/architecture-guide/bdm-guide/index.md +++ b/src/content/docs/tutorial/onboarding/architecture-guide/bdm-guide/index.md @@ -1,6 +1,7 @@ --- title: Business Data Model Guide slug: 'tutorial/onboarding/architecture-guide/bdm-guide' +description: 'Design a Business Data Model — a shared, business-oriented view of your data that guides how projects and pipelines are structured.' --- diff --git a/src/content/docs/tutorial/onboarding/architecture-guide/index.md b/src/content/docs/tutorial/onboarding/architecture-guide/index.md index b48670c92..0a37bf8a6 100644 --- a/src/content/docs/tutorial/onboarding/architecture-guide/index.md +++ b/src/content/docs/tutorial/onboarding/architecture-guide/index.md @@ -1,6 +1,7 @@ --- title: Multi-Project Architecture Guide slug: 'tutorial/onboarding/architecture-guide' +description: 'How to organize Keboola projects — single project, multi-project architectures, and proven setups for organizations of different sizes.' --- diff --git a/src/content/docs/tutorial/onboarding/cheat-sheet/index.md b/src/content/docs/tutorial/onboarding/cheat-sheet/index.md index 230e19ee9..ded2c3534 100644 --- a/src/content/docs/tutorial/onboarding/cheat-sheet/index.md +++ b/src/content/docs/tutorial/onboarding/cheat-sheet/index.md @@ -1,6 +1,7 @@ --- title: "Cheat Sheet: Best Practices" slug: 'tutorial/onboarding/cheat-sheet' +description: 'Best practices for everyday Keboola work — credentials, incremental loads, transformations, flows, error handling, and development branches.' --- How you use Keboola can differ significantly. From setting up automated data pipelines with Flows to managing components such as data sources, destinations, diff --git a/src/content/docs/tutorial/onboarding/governance-guide/index.md b/src/content/docs/tutorial/onboarding/governance-guide/index.md index 5763189f5..9fd50d29a 100644 --- a/src/content/docs/tutorial/onboarding/governance-guide/index.md +++ b/src/content/docs/tutorial/onboarding/governance-guide/index.md @@ -1,6 +1,7 @@ --- title: Keboola Governance Guide slug: 'tutorial/onboarding/governance-guide' +description: 'Governance practices for Keboola — access management, security, monitoring, and operational responsibilities across your projects.' --- Welcome to the Keboola Governance Guide! Here, we cover everything you need to know about managing and monitoring your use of our platform, diff --git a/src/content/docs/tutorial/onboarding/index.md b/src/content/docs/tutorial/onboarding/index.md index f986eb242..b5533a1a2 100644 --- a/src/content/docs/tutorial/onboarding/index.md +++ b/src/content/docs/tutorial/onboarding/index.md @@ -1,6 +1,7 @@ --- title: Keboola Platform Onboarding slug: 'tutorial/onboarding' +description: 'Start here after your proof of concept — a guided path through architecture planning, governance, usage documentation, and best practices.' --- diff --git a/src/content/docs/tutorial/onboarding/usage-blueprint/index.md b/src/content/docs/tutorial/onboarding/usage-blueprint/index.md index d2206a2f6..bd3aa4b10 100644 --- a/src/content/docs/tutorial/onboarding/usage-blueprint/index.md +++ b/src/content/docs/tutorial/onboarding/usage-blueprint/index.md @@ -1,6 +1,7 @@ --- title: Keboola Platform Usage Blueprint slug: 'tutorial/onboarding/usage-blueprint' +description: 'A customizable template for documenting how your organization uses Keboola — naming conventions, responsibilities, and operating rules.' --- > Welcome to your personalized Keboola Platform Usage Blueprint Document! diff --git a/src/content/docs/tutorial/write/index.md b/src/content/docs/tutorial/write/index.md index cbe57c809..4d43b7c40 100644 --- a/src/content/docs/tutorial/write/index.md +++ b/src/content/docs/tutorial/write/index.md @@ -1,6 +1,7 @@ --- title: "Part 3: Writing Data" slug: 'tutorial/write' +description: 'Part 3 of the tutorial — write the transformed table from Storage to Google Sheets with a data destination connector (reverse ETL).' ---