Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
160 changes: 160 additions & 0 deletions content/de/developer/integration/big-data/clickhouse.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,160 @@
---
title: "ClickHouse"
description: "Run ClickHouse with an S3 disk backed by RustFS for MergeTree table data."
---

This guide connects [ClickHouse](https://github.com/ClickHouse/ClickHouse) — the real-time OLAP database — to **RustFS** through ClickHouse's S3 disk storage policy. You will start a ClickHouse server with Docker, create a MergeTree table that stores its parts on RustFS, insert rows, and verify that the table data lives in the bucket. The workflow was verified with `clickhouse/clickhouse-server:25.8` and `rustfs/rustfs-x86-musl:v2.3.1`.

You need Docker. This deployment is intended for local integration testing, not production.

## Architecture

```mermaid
flowchart LR
Client["SQL client"] -->|"queries"| CH["ClickHouse :8123"]
CH -->|"MergeTree parts"| RustFS["RustFS :9000"]
```

The `rustfs` disk is a ClickHouse S3 disk pointed at the `clickhouse-data` bucket. Tables created with the matching storage policy write their parts — data, index, and checksum files — to the bucket instead of the local filesystem.

## 1. Create the project files

Create the bucket first — ClickHouse does not create buckets:

```bash
rc alias set rustfs http://<your-rustfs-endpoint>:9000 <your-access-key> <your-secret-key>
rc mb rustfs/clickhouse-data
```

Create the storage configuration, replacing both credential placeholders:

```xml title="storage.xml"
<clickhouse>
<storage_configuration>
<disks>
<rustfs>
<type>s3</type>
<endpoint>http://rustfs:9000/clickhouse-data/</endpoint>
<access_key_id><your-access-key></access_key_id>
<secret_access_key><your-secret-key></secret_access_key>
</rustfs>
</disks>
<policies>
<rustfs_policy>
<volumes>
<main>
<disk>rustfs</disk>
</main>
</volumes>
</rustfs_policy>
</policies>
</storage_configuration>
</clickhouse>
```

The endpoint must end with `/` and includes the bucket name as the first path segment. Inside the Compose network the hostname is `rustfs`; from the host use `http://localhost:9000/clickhouse-data/`.

Start ClickHouse with the configuration mounted:

```bash
docker run -d --name clickhouse --network oo-rustfs_default \
-p 8123:8123 \
-e CLICKHOUSE_PASSWORD=<your-clickhouse-password> \
-v "$PWD/storage.xml":/etc/clickhouse-server/config.d/storage.xml:ro \
clickhouse/clickhouse-server:25.8
```

## 2. Create a table on the S3 disk

Wait for the HTTP interface, then create a database and a MergeTree table with the storage policy:

```bash
curl "http://localhost:8123/?password=<your-clickhouse-password>" \
--data-binary "CREATE DATABASE rustfs_demo"

curl "http://localhost:8123/?password=<your-clickhouse-password>" \
--data-binary "CREATE TABLE rustfs_demo.events
(id UInt32, name String)
ENGINE = MergeTree ORDER BY id
SETTINGS storage_policy = 'rustfs_policy'"

curl "http://localhost:8123/?password=<your-clickhouse-password>" \
--data-binary "INSERT INTO rustfs_demo.events
VALUES (1, 'clickhouse-on-rustfs'), (2, 'second')"
```

Read the rows back and confirm ClickHouse reports the part on the `rustfs` disk:

```bash
curl "http://localhost:8123/?password=<your-clickhouse-password>" \
--data-binary "SELECT count(), any(name) FROM rustfs_demo.events"

curl "http://localhost:8123/?password=<your-clickhouse-password>" \
--data-binary "SELECT name, disk_name FROM system.parts
WHERE database = 'rustfs_demo' AND active"
```

```text
2 clickhouse-on-rustfs
all_1_1_0 rustfs
```

## 3. Verify objects in RustFS

List the bucket:

```bash
rc ls rustfs/clickhouse-data/ -r
```

ClickHouse writes each part as content-addressed blobs. The output contains several small objects, and the count grows as more parts are written:

```text
dtg/hpsyncexixvdnsgorvseobogcgowg
dzp/zfblobhsatzdveqdrsfqcupkdehja
izg/gvhchqobrizpkdftuvlmakkoutwps
```

![ClickHouse parts stored in the RustFS Console](./images/rustfs-clickhouse-disk.png)

The data survives a container restart because the parts live in RustFS:

```bash
docker restart clickhouse
curl "http://localhost:8123/?password=<your-clickhouse-password>" \
--data-binary "SELECT count() FROM rustfs_demo.events"
```

## 4. Stop or reset the deployment

Stop the server while keeping the data:

```bash
docker rm -f clickhouse
```

The parts stay in the `clickhouse-data` bucket and the table can be queried again after the next start. To delete the data, remove the bucket:

```bash
rc rb rustfs/clickhouse-data --force
```

## Troubleshooting

### `REQUIRED_PASSWORD` on every query

ClickHouse 25.8 images require a password for the `default` user. Set `CLICKHOUSE_PASSWORD` on the container and pass the same value as the `password` query parameter, as shown above.

### Table creation fails with a disk or endpoint error

Confirm that the bucket exists before the table is created, that the endpoint ends with `/`, and that the credentials match the RustFS deployment. Check the server log for the underlying S3 error:

```bash
docker logs clickhouse | grep -i s3 | tail
```

## Next steps

- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional ClickHouse operations.
- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token).
- Follow the [ClickHouse S3 disk documentation](https://clickhouse.com/docs/engines/table-engines/mergetree-family/mergetree#table_engine-mergetree-s3) to add a cache disk or a tiered hot/cold policy.
161 changes: 161 additions & 0 deletions content/de/developer/integration/big-data/doris.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,161 @@
---
title: "Apache Doris"
description: "Back up Apache Doris tables to RustFS through an S3 repository and restore them."
---

This guide connects [Apache Doris](https://github.com/apache/doris) — the real-time analytical data warehouse — to **RustFS** through an S3 backup repository. You will start an all-in-one Doris container, create an S3 repository pointing at a RustFS bucket, back up a table, drop it, and restore it from RustFS. The workflow was verified with `apache/doris:all-in-one-4.1.3` and `rustfs/rustfs-x86-musl:v2.3.1`.

You need Docker. This deployment is intended for local integration testing, not production.

## Architecture

```mermaid
flowchart LR
Client["SQL client"] -->|"queries"| Doris["Doris FE/BE"]
Doris -->|"BACKUP / RESTORE"| RustFS["RustFS :9000"]
```

The repository is a named S3 location under the `doris-backups` bucket. `BACKUP SNAPSHOT` uploads table metadata and tablet data files; `RESTORE SNAPSHOT` downloads them into a new table.

## 1. Start Doris and create the repository

Create the bucket first — Doris does not create buckets:

```bash
rc alias set rustfs http://<your-rustfs-endpoint>:9000 <your-access-key> <your-secret-key>
rc mb rustfs/doris-backups
```

Start the all-in-one container on the same Docker network as RustFS:

```bash
docker run -d --name doris --network oo-rustfs_default \
-p 8030:8030 -p 9030:9030 apache/doris:all-in-one-4.1.3
```

Wait for the frontend to become healthy, then connect with the MySQL protocol (port 9030, user `root`, no password in the all-in-one image).

Create a test table with rows:

```sql
CREATE DATABASE rustfs_demo;
CREATE TABLE rustfs_demo.events
(id INT, name VARCHAR(50))
DISTRIBUTED BY HASH(id) BUCKETS 1
PROPERTIES ("replication_num" = "1");
INSERT INTO rustfs_demo.events VALUES (1, 'doris-on-rustfs'), (2, 'backup-test');
```

Create the S3 repository, replacing the endpoint with the IP address of the RustFS container and both credential placeholders:

```sql
CREATE REPOSITORY `rustfs_repo`
WITH S3
ON LOCATION "s3://doris-backups/rustfs-repo"
PROPERTIES (
"AWS_ENDPOINT" = "http://<rustfs-container-ip>:9000",
"AWS_ACCESS_KEY" = "<your-access-key>",
"AWS_SECRET_KEY" = "<your-secret-key>",
"AWS_REGION" = "us-east-1",
"AWS_PATH_STYLE_ACCESS" = "true"
);
```

Doris 4.1 resolves the bucket into the endpoint hostname even with `AWS_PATH_STYLE_ACCESS` enabled, so a hostname endpoint fails with `UnknownHostException: doris-backups.rustfs`. Using the container IP address forces path-style requests and works; `SHOW REPOSITORIES` confirms the repository registered with an empty `ErrMsg`.

## 2. Back up a table to RustFS

Take a snapshot of the table:

```sql
BACKUP SNAPSHOT rustfs_demo.demo_snapshot
TO rustfs_repo
ON (events);
```

The statement returns immediately; the backup job runs in the background. Watch its state:

```sql
SHOW BACKUP;
```

Wait until `State` reaches `FINISHED` — the snapshot metadata and tablet data files are now objects in the bucket.

## 3. Verify the backup in RustFS

List the bucket:

```bash
rc ls rustfs/doris-backups/ -r
```

The repository stores a repository descriptor, the snapshot metadata, and the tablet files:

```text
rustfs-repo/__palo_repository_rustfs_repo/__repo_info
rustfs-repo/__palo_repository_rustfs_repo/__ss_demo_snapshot/__meta.d50ecf9b...
rustfs-repo/__palo_repository_rustfs_repo/__ss_demo_snapshot/__ss_content/.../...dat...
```

![Doris backup objects stored in the RustFS Console](./images/rustfs-doris-backup.png)

## 4. Restore the table from RustFS

Drop the table and restore it from the snapshot. The timestamp comes from the snapshot name shown by `SHOW SNAPSHOT ON REPOSITORY rustfs_repo;`:

```sql
DROP TABLE rustfs_demo.events;

RESTORE SNAPSHOT rustfs_demo.demo_snapshot
FROM rustfs_repo
ON (events)
PROPERTIES (
"backup_timestamp" = "2026-09-21-16-35-12",
"replication_num" = "1"
);
```

Wait for the restore job to finish and confirm the data:

```sql
SHOW RESTORE;
SELECT count(*) FROM rustfs_demo.events;
```

```text
2
```

## 5. Stop or reset the deployment

Stop Doris while keeping the data:

```bash
docker rm -f doris
```

The backup stays in the `doris-backups` bucket and can be restored into any Doris cluster that registers the same repository. To delete it, remove the bucket:

```bash
rc rb rustfs/doris-backups --force
```

## Troubleshooting

### `UnknownHostException: doris-backups.rustfs` when creating the repository

Doris is building a virtual-hosted hostname from the bucket and endpoint. Use the RustFS container IP address in `AWS_ENDPOINT` together with `AWS_PATH_STYLE_ACCESS = "true"`, as shown above.

### The backup stays in `SNAPSHOTING` for a long time

The backend uploads the tablet files. Confirm the backend is healthy (`SHOW BACKENDS;`) and can reach the endpoint; the all-in-one image needs a minute or two after start before both processes report ready.

### `Failed to create repository: ... file status`

The bucket does not exist or the credentials are wrong. Create `doris-backups` with `rc mb` and re-check the access key pair.

## Next steps

- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Doris operations.
- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token).
- Follow the [Doris backup and restore documentation](https://doris.apache.org/docs/data-operate/backup-restore/) to schedule periodic snapshots.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
3 changes: 3 additions & 0 deletions content/de/developer/integration/big-data/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,11 +7,14 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo

## Systems

- [ClickHouse](./clickhouse.md)
- [Iceberg](./iceberg.md)
- [PyIceberg](./pyiceberg.md)
- [Milvus](./milvus.md)
- [MLflow](./mlflow.md)
- [OpenDAL](./opendal.md)
- [DuckDB](./duckdb.md)
- [Doris](./doris.md)
- [InfluxDB](./influxdb.md)
- [Spark](./spark.md)
- [Flink](./flink.md)
Expand Down
6 changes: 5 additions & 1 deletion content/de/developer/integration/big-data/meta.json
Original file line number Diff line number Diff line change
@@ -1,14 +1,18 @@
{
"title": "Data Analytics",
"pages": [
"clickhouse",
"iceberg",
"pyiceberg",
"milvus",
"mlflow",
"opendal",
"duckdb",
"doris",
"influxdb",
"spark",
"flink",
"trino"
"trino",
"zeppelin"
]
}
Loading
Loading