diff --git a/content/de/developer/integration/big-data/clickhouse.md b/content/de/developer/integration/big-data/clickhouse.md new file mode 100644 index 00000000..ac9d8d4b --- /dev/null +++ b/content/de/developer/integration/big-data/clickhouse.md @@ -0,0 +1,160 @@ +--- +title: "ClickHouse" +description: "Run ClickHouse with an S3 disk backed by RustFS for MergeTree table data." +--- + +This guide connects [ClickHouse](https://github.com/ClickHouse/ClickHouse) — the real-time OLAP database — to **RustFS** through ClickHouse's S3 disk storage policy. You will start a ClickHouse server with Docker, create a MergeTree table that stores its parts on RustFS, insert rows, and verify that the table data lives in the bucket. The workflow was verified with `clickhouse/clickhouse-server:25.8` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["SQL client"] -->|"queries"| CH["ClickHouse :8123"] + CH -->|"MergeTree parts"| RustFS["RustFS :9000"] +``` + +The `rustfs` disk is a ClickHouse S3 disk pointed at the `clickhouse-data` bucket. Tables created with the matching storage policy write their parts — data, index, and checksum files — to the bucket instead of the local filesystem. + +## 1. Create the project files + +Create the bucket first — ClickHouse does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/clickhouse-data +``` + +Create the storage configuration, replacing both credential placeholders: + +```xml title="storage.xml" + + + + + s3 + http://rustfs:9000/clickhouse-data/ + + + + + + + +
+ rustfs +
+
+
+
+
+
+``` + +The endpoint must end with `/` and includes the bucket name as the first path segment. Inside the Compose network the hostname is `rustfs`; from the host use `http://localhost:9000/clickhouse-data/`. + +Start ClickHouse with the configuration mounted: + +```bash +docker run -d --name clickhouse --network oo-rustfs_default \ + -p 8123:8123 \ + -e CLICKHOUSE_PASSWORD= \ + -v "$PWD/storage.xml":/etc/clickhouse-server/config.d/storage.xml:ro \ + clickhouse/clickhouse-server:25.8 +``` + +## 2. Create a table on the S3 disk + +Wait for the HTTP interface, then create a database and a MergeTree table with the storage policy: + +```bash +curl "http://localhost:8123/?password=" \ + --data-binary "CREATE DATABASE rustfs_demo" + +curl "http://localhost:8123/?password=" \ + --data-binary "CREATE TABLE rustfs_demo.events + (id UInt32, name String) + ENGINE = MergeTree ORDER BY id + SETTINGS storage_policy = 'rustfs_policy'" + +curl "http://localhost:8123/?password=" \ + --data-binary "INSERT INTO rustfs_demo.events + VALUES (1, 'clickhouse-on-rustfs'), (2, 'second')" +``` + +Read the rows back and confirm ClickHouse reports the part on the `rustfs` disk: + +```bash +curl "http://localhost:8123/?password=" \ + --data-binary "SELECT count(), any(name) FROM rustfs_demo.events" + +curl "http://localhost:8123/?password=" \ + --data-binary "SELECT name, disk_name FROM system.parts + WHERE database = 'rustfs_demo' AND active" +``` + +```text +2 clickhouse-on-rustfs +all_1_1_0 rustfs +``` + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/clickhouse-data/ -r +``` + +ClickHouse writes each part as content-addressed blobs. The output contains several small objects, and the count grows as more parts are written: + +```text +dtg/hpsyncexixvdnsgorvseobogcgowg +dzp/zfblobhsatzdveqdrsfqcupkdehja +izg/gvhchqobrizpkdftuvlmakkoutwps +``` + +![ClickHouse parts stored in the RustFS Console](./images/rustfs-clickhouse-disk.png) + +The data survives a container restart because the parts live in RustFS: + +```bash +docker restart clickhouse +curl "http://localhost:8123/?password=" \ + --data-binary "SELECT count() FROM rustfs_demo.events" +``` + +## 4. Stop or reset the deployment + +Stop the server while keeping the data: + +```bash +docker rm -f clickhouse +``` + +The parts stay in the `clickhouse-data` bucket and the table can be queried again after the next start. To delete the data, remove the bucket: + +```bash +rc rb rustfs/clickhouse-data --force +``` + +## Troubleshooting + +### `REQUIRED_PASSWORD` on every query + +ClickHouse 25.8 images require a password for the `default` user. Set `CLICKHOUSE_PASSWORD` on the container and pass the same value as the `password` query parameter, as shown above. + +### Table creation fails with a disk or endpoint error + +Confirm that the bucket exists before the table is created, that the endpoint ends with `/`, and that the credentials match the RustFS deployment. Check the server log for the underlying S3 error: + +```bash +docker logs clickhouse | grep -i s3 | tail +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional ClickHouse operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [ClickHouse S3 disk documentation](https://clickhouse.com/docs/engines/table-engines/mergetree-family/mergetree#table_engine-mergetree-s3) to add a cache disk or a tiered hot/cold policy. diff --git a/content/de/developer/integration/big-data/doris.md b/content/de/developer/integration/big-data/doris.md new file mode 100644 index 00000000..1c3665b5 --- /dev/null +++ b/content/de/developer/integration/big-data/doris.md @@ -0,0 +1,161 @@ +--- +title: "Apache Doris" +description: "Back up Apache Doris tables to RustFS through an S3 repository and restore them." +--- + +This guide connects [Apache Doris](https://github.com/apache/doris) — the real-time analytical data warehouse — to **RustFS** through an S3 backup repository. You will start an all-in-one Doris container, create an S3 repository pointing at a RustFS bucket, back up a table, drop it, and restore it from RustFS. The workflow was verified with `apache/doris:all-in-one-4.1.3` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["SQL client"] -->|"queries"| Doris["Doris FE/BE"] + Doris -->|"BACKUP / RESTORE"| RustFS["RustFS :9000"] +``` + +The repository is a named S3 location under the `doris-backups` bucket. `BACKUP SNAPSHOT` uploads table metadata and tablet data files; `RESTORE SNAPSHOT` downloads them into a new table. + +## 1. Start Doris and create the repository + +Create the bucket first — Doris does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/doris-backups +``` + +Start the all-in-one container on the same Docker network as RustFS: + +```bash +docker run -d --name doris --network oo-rustfs_default \ + -p 8030:8030 -p 9030:9030 apache/doris:all-in-one-4.1.3 +``` + +Wait for the frontend to become healthy, then connect with the MySQL protocol (port 9030, user `root`, no password in the all-in-one image). + +Create a test table with rows: + +```sql +CREATE DATABASE rustfs_demo; +CREATE TABLE rustfs_demo.events + (id INT, name VARCHAR(50)) + DISTRIBUTED BY HASH(id) BUCKETS 1 + PROPERTIES ("replication_num" = "1"); +INSERT INTO rustfs_demo.events VALUES (1, 'doris-on-rustfs'), (2, 'backup-test'); +``` + +Create the S3 repository, replacing the endpoint with the IP address of the RustFS container and both credential placeholders: + +```sql +CREATE REPOSITORY `rustfs_repo` + WITH S3 + ON LOCATION "s3://doris-backups/rustfs-repo" + PROPERTIES ( + "AWS_ENDPOINT" = "http://:9000", + "AWS_ACCESS_KEY" = "", + "AWS_SECRET_KEY" = "", + "AWS_REGION" = "us-east-1", + "AWS_PATH_STYLE_ACCESS" = "true" + ); +``` + +Doris 4.1 resolves the bucket into the endpoint hostname even with `AWS_PATH_STYLE_ACCESS` enabled, so a hostname endpoint fails with `UnknownHostException: doris-backups.rustfs`. Using the container IP address forces path-style requests and works; `SHOW REPOSITORIES` confirms the repository registered with an empty `ErrMsg`. + +## 2. Back up a table to RustFS + +Take a snapshot of the table: + +```sql +BACKUP SNAPSHOT rustfs_demo.demo_snapshot + TO rustfs_repo + ON (events); +``` + +The statement returns immediately; the backup job runs in the background. Watch its state: + +```sql +SHOW BACKUP; +``` + +Wait until `State` reaches `FINISHED` — the snapshot metadata and tablet data files are now objects in the bucket. + +## 3. Verify the backup in RustFS + +List the bucket: + +```bash +rc ls rustfs/doris-backups/ -r +``` + +The repository stores a repository descriptor, the snapshot metadata, and the tablet files: + +```text +rustfs-repo/__palo_repository_rustfs_repo/__repo_info +rustfs-repo/__palo_repository_rustfs_repo/__ss_demo_snapshot/__meta.d50ecf9b... +rustfs-repo/__palo_repository_rustfs_repo/__ss_demo_snapshot/__ss_content/.../...dat... +``` + +![Doris backup objects stored in the RustFS Console](./images/rustfs-doris-backup.png) + +## 4. Restore the table from RustFS + +Drop the table and restore it from the snapshot. The timestamp comes from the snapshot name shown by `SHOW SNAPSHOT ON REPOSITORY rustfs_repo;`: + +```sql +DROP TABLE rustfs_demo.events; + +RESTORE SNAPSHOT rustfs_demo.demo_snapshot + FROM rustfs_repo + ON (events) + PROPERTIES ( + "backup_timestamp" = "2026-09-21-16-35-12", + "replication_num" = "1" + ); +``` + +Wait for the restore job to finish and confirm the data: + +```sql +SHOW RESTORE; +SELECT count(*) FROM rustfs_demo.events; +``` + +```text +2 +``` + +## 5. Stop or reset the deployment + +Stop Doris while keeping the data: + +```bash +docker rm -f doris +``` + +The backup stays in the `doris-backups` bucket and can be restored into any Doris cluster that registers the same repository. To delete it, remove the bucket: + +```bash +rc rb rustfs/doris-backups --force +``` + +## Troubleshooting + +### `UnknownHostException: doris-backups.rustfs` when creating the repository + +Doris is building a virtual-hosted hostname from the bucket and endpoint. Use the RustFS container IP address in `AWS_ENDPOINT` together with `AWS_PATH_STYLE_ACCESS = "true"`, as shown above. + +### The backup stays in `SNAPSHOTING` for a long time + +The backend uploads the tablet files. Confirm the backend is healthy (`SHOW BACKENDS;`) and can reach the endpoint; the all-in-one image needs a minute or two after start before both processes report ready. + +### `Failed to create repository: ... file status` + +The bucket does not exist or the credentials are wrong. Create `doris-backups` with `rc mb` and re-check the access key pair. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Doris operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Doris backup and restore documentation](https://doris.apache.org/docs/data-operate/backup-restore/) to schedule periodic snapshots. diff --git a/content/de/developer/integration/big-data/images/rustfs-clickhouse-disk.png b/content/de/developer/integration/big-data/images/rustfs-clickhouse-disk.png new file mode 100644 index 00000000..9225bf66 Binary files /dev/null and b/content/de/developer/integration/big-data/images/rustfs-clickhouse-disk.png differ diff --git a/content/de/developer/integration/big-data/images/rustfs-doris-backup.png b/content/de/developer/integration/big-data/images/rustfs-doris-backup.png new file mode 100644 index 00000000..9a936276 Binary files /dev/null and b/content/de/developer/integration/big-data/images/rustfs-doris-backup.png differ diff --git a/content/de/developer/integration/big-data/images/rustfs-opendal-objects.png b/content/de/developer/integration/big-data/images/rustfs-opendal-objects.png new file mode 100644 index 00000000..d7129423 Binary files /dev/null and b/content/de/developer/integration/big-data/images/rustfs-opendal-objects.png differ diff --git a/content/de/developer/integration/big-data/images/rustfs-zeppelin-notebook.png b/content/de/developer/integration/big-data/images/rustfs-zeppelin-notebook.png new file mode 100644 index 00000000..9c5fe782 Binary files /dev/null and b/content/de/developer/integration/big-data/images/rustfs-zeppelin-notebook.png differ diff --git a/content/de/developer/integration/big-data/index.md b/content/de/developer/integration/big-data/index.md index ef0b7cba..342e0be6 100644 --- a/content/de/developer/integration/big-data/index.md +++ b/content/de/developer/integration/big-data/index.md @@ -7,11 +7,14 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo ## Systems +- [ClickHouse](./clickhouse.md) - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) - [Milvus](./milvus.md) - [MLflow](./mlflow.md) +- [OpenDAL](./opendal.md) - [DuckDB](./duckdb.md) +- [Doris](./doris.md) - [InfluxDB](./influxdb.md) - [Spark](./spark.md) - [Flink](./flink.md) diff --git a/content/de/developer/integration/big-data/meta.json b/content/de/developer/integration/big-data/meta.json index 87e29f91..ee7d69b5 100644 --- a/content/de/developer/integration/big-data/meta.json +++ b/content/de/developer/integration/big-data/meta.json @@ -1,14 +1,18 @@ { "title": "Data Analytics", "pages": [ + "clickhouse", "iceberg", "pyiceberg", "milvus", "mlflow", + "opendal", "duckdb", + "doris", "influxdb", "spark", "flink", - "trino" + "trino", + "zeppelin" ] } diff --git a/content/de/developer/integration/big-data/opendal.md b/content/de/developer/integration/big-data/opendal.md new file mode 100644 index 00000000..074b7475 --- /dev/null +++ b/content/de/developer/integration/big-data/opendal.md @@ -0,0 +1,108 @@ +--- +title: "OpenDAL" +description: "Access RustFS objects from applications through the Apache OpenDAL data access layer." +--- + +This guide connects [Apache OpenDAL](https://github.com/apache/opendal) — the unified data access layer — to **RustFS** through its `s3` service. You will run the OpenDAL Python binding against RustFS, write and read an object, list a prefix, and delete it. The workflow was verified with the `opendal` Python package 0.46 on `python:3.12-slim` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and Python 3.9 or later. OpenDAL supports the same `s3` service from Rust, Java, Node.js, and Go bindings with equivalent settings. + +## Architecture + +```mermaid +flowchart LR + App["Application"] -->|"Operator API"| OpenDAL["OpenDAL"] + OpenDAL -->|"s3 service"| RustFS["RustFS :9000"] +``` + +An OpenDAL `Operator` bound to the `s3` service exposes one uniform API — `write`, `read`, `stat`, `list`, `delete` — over the bucket, so the same code runs against S3, RustFS, or any other supported service by changing the connection settings. + +## 1. Set up the project + +Install the Python binding: + +```bash +pip install opendal +``` + +Create the script, replacing all connection placeholders: + +```python title="opendal_demo.py" +import opendal + +op = opendal.Operator( + "s3", + endpoint="http://:9000", + bucket="my-bucket", + access_key_id="", + secret_access_key="", + region="us-east-1", +) +op.write("opendal-demo/hello.txt", b"hello from opendal against rustfs") +print("read-back:", op.read("opendal-demo/hello.txt")) +print("content_length:", op.stat("opendal-demo/hello.txt").content_length) +for entry in op.list("opendal-demo/"): + print("listed:", entry.path) +op.delete("opendal-demo/hello.txt") +print("deleted:", not op.exists("opendal-demo/hello.txt")) +``` + +The credential option names are `access_key_id` and `secret_access_key` — the shorter `access_key` names do not exist and fail with a signing error. Path-style addressing is the default for non-AWS endpoints. + +## 2. Run the demo + +Run the script from a machine that can reach RustFS: + +```bash +python opendal_demo.py +``` + +```text +read-back: b"hello from opendal against rustfs" +content_length: 33 +listed: opendal-demo/hello.txt +deleted: True +``` + +The round trip exercises the full object lifecycle: write uploads the bytes, `read` fetches them back, `stat` returns the object size, `list` enumerates the prefix, and `delete` removes the object. + +## 3. Verify objects in RustFS + +Comment out the final `op.delete` line, run the script again, and list the prefix in RustFS: + +```bash +rc ls rustfs/my-bucket/opendal-demo/ -r +``` + +```text +hello.txt +data/rows.csv +``` + +![OpenDAL objects stored in the RustFS Console](./images/rustfs-opendal-objects.png) + +The objects visible in the RustFS Console are exactly the paths the OpenDAL API wrote. + +## 4. Stop or reset + +OpenDAL is a library and holds no state of its own. To clean up the demo objects: + +```bash +rc rm rustfs/my-bucket/opendal-demo/ --recursive --force +``` + +## Troubleshooting + +### `failed to load signing credential` + +The operator received no usable credentials. Use the exact option names `access_key_id` and `secret_access_key`; other spellings are silently ignored and signing then fails. + +### Connection or DNS errors on write + +Confirm the endpoint includes the scheme and port and is reachable from the application. Inside a Compose network the hostname is `rustfs`; from the host use `http://localhost:9000`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional OpenDAL operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [OpenDAL documentation](https://opendal.apache.org/docs/) to use the same operator from Rust, Java, or Node.js. diff --git a/content/de/developer/integration/big-data/zeppelin.md b/content/de/developer/integration/big-data/zeppelin.md new file mode 100644 index 00000000..e7f47375 --- /dev/null +++ b/content/de/developer/integration/big-data/zeppelin.md @@ -0,0 +1,109 @@ +--- +title: "Apache Zeppelin" +description: "Store Apache Zeppelin notebooks in RustFS through the S3 notebook repository." +--- + +This guide connects [Apache Zeppelin](https://github.com/apache/zeppelin) — the web-based notebook for data analytics — to **RustFS** through Zeppelin's S3 notebook storage. You will start Zeppelin pointed at a RustFS bucket, create a note, and verify that the notebook file is stored in RustFS and survives a restart. The workflow was verified with `apache/zeppelin:0.12.0` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Browser["Browser"] -->|"notebook edits"| Z["Zeppelin :8080"] + Z -->|".zpln files"| RustFS["RustFS :9000"] +``` + +With the `S3NotebookRepo` storage class, every note is persisted as a `.zpln` JSON file under the `user/notebook/` prefix of the bucket. Zeppelin reads and writes the bucket directly, so notes survive container restarts and can be shared between instances. + +## 1. Start Zeppelin + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +AWS_ACCESS_KEY_ID= +AWS_SECRET_ACCESS_KEY= +``` + +Start Zeppelin on the same Docker network as RustFS with the S3 storage settings: + +```bash +docker run -d --name zeppelin --network oo-rustfs_default \ + -p 8080:8080 \ + -e AWS_ACCESS_KEY_ID \ + -e AWS_SECRET_ACCESS_KEY \ + -e ZEPPELIN_NOTEBOOK_STORAGE=org.apache.zeppelin.notebook.repo.S3NotebookRepo \ + -e ZEPPELIN_NOTEBOOK_S3_BUCKET=my-bucket \ + -e ZEPPELIN_NOTEBOOK_S3_ENDPOINT=http://rustfs:9000 \ + -e ZEPPELIN_NOTEBOOK_S3_PATH_STYLE_ACCESS=true \ + apache/zeppelin:0.12.0 +``` + +Zeppelin reads `ZEPPELIN_*` environment variables as configuration properties, so no `zeppelin-site.xml` edit is needed. This guide uses the existing `my-bucket`; notes land under its `user/notebook/` prefix, which the S3 storage creates on demand. + +## 2. Create a note + +Wait for the UI on `http://localhost:8080`, then create a note named `rustfs-demo` in the notebook list and add a paragraph, or use the REST API: + +```bash +NOTE=$(curl -s -X POST "http://localhost:8080/api/notebook" \ + -H "Content-Type: application/json" \ + -d '{"name": "rustfs-demo"}' | python3 -c "import json,sys; print(json.load(sys.stdin)['body'])") +echo "note id: $NOTE" +``` + +## 3. Verify the notebook in RustFS + +List the notebook prefix: + +```bash +rc ls rustfs/my-bucket/user/notebook/ -r +``` + +The note is stored as a JSON file named after the note and its ID: + +```text +user/notebook/rustfs-demo_2N4PY7UY5.zpln +``` + +![Zeppelin notebooks stored in the RustFS Console](./images/rustfs-zeppelin-notebook.png) + +Notes survive a restart because Zeppelin loads them from the bucket: + +```bash +docker restart zeppelin +curl -s "http://localhost:8080/api/notebook" | head -c 200 +``` + +The note list contains `2N4PY7UY5` again after the restart. + +## 4. Stop or reset the deployment + +Stop Zeppelin while keeping the notes: + +```bash +docker rm -f zeppelin +``` + +The notes stay in the `user/notebook/` prefix of `my-bucket`. To delete them, remove the prefix: + +```bash +rc rm rustfs/my-bucket/user/notebook/ --recursive --force +``` + +## Troubleshooting + +### Zeppelin starts but notes never appear in the bucket + +Confirm the three `ZEPPELIN_NOTEBOOK_S3_*` variables are set and that the credentials environment variables reach the container — the S3 repository is initialized at startup, so the container must be recreated after any change. + +### `UnknownHostException: my-bucket.rustfs` + +The path-style flag was not picked up. Keep `ZEPPELIN_NOTEBOOK_S3_PATH_STYLE_ACCESS=true` exactly as shown; with path-style disabled Zeppelin treats the bucket as a hostname. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Zeppelin operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Zeppelin storage documentation](https://zeppelin.apache.org/docs/latest/setup/storage/storage.html#notebook-storage-in-s3) to organize notes in per-user prefixes. diff --git a/content/de/developer/integration/devops/images/rustfs-jenkins-artifacts.png b/content/de/developer/integration/devops/images/rustfs-jenkins-artifacts.png new file mode 100644 index 00000000..19c8cf87 Binary files /dev/null and b/content/de/developer/integration/devops/images/rustfs-jenkins-artifacts.png differ diff --git a/content/de/developer/integration/devops/index.md b/content/de/developer/integration/devops/index.md index 58fc491d..86b751d4 100644 --- a/content/de/developer/integration/devops/index.md +++ b/content/de/developer/integration/devops/index.md @@ -9,6 +9,7 @@ Nutzen Sie **RustFS** als Objektspeicher-Layer für DevOps-Plattformen und Infra - [Elasticsearch](./elasticsearch.md) - [Gitea](./gitea.md) +- [Jenkins](./jenkins.md) - [Terraform](./terraform.md) Speichern Sie Artefakte, State und Telemetriedaten in dedizierten Buckets und beschränken Sie die Anmeldeinformationen auf die erforderlichen Bucket-Operationen. diff --git a/content/de/developer/integration/devops/jenkins.md b/content/de/developer/integration/devops/jenkins.md new file mode 100644 index 00000000..1d78f31b --- /dev/null +++ b/content/de/developer/integration/devops/jenkins.md @@ -0,0 +1,144 @@ +--- +title: "Jenkins" +description: "Store Jenkins build artifacts in RustFS with the Artifact Manager on S3 plugin." +--- + +This guide connects [Jenkins](https://github.com/jenkinsci/jenkins) — the automation server — to **RustFS** through the Artifact Manager on S3 plugin. You will start Jenkins with the plugin, point its artifact manager at a RustFS bucket, run a job that archives an artifact, and verify that the artifact is stored in RustFS. The workflow was verified with `jenkins/jenkins:lts-jdk17` (Jenkins 2.5xx LTS) and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Dev["Developer"] -->|"trigger build"| J["Jenkins :8080"] + J -->|"archiveArtifacts"| RustFS["RustFS :9000"] +``` + +With the plugin active, every job that publishes artifacts through the standard `archiveArtifacts` step (or `stash`/`unstash`) uploads them to the `jenkins-artifacts` bucket under the configured prefix instead of storing them on the controller disk. + +## 1. Create the project files + +Create the bucket first — the plugin validates the bucket but does not create it: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/jenkins-artifacts +``` + +Create a Dockerfile that installs the plugin into the Jenkins LTS image: + +```dockerfile title="Dockerfile" +FROM jenkins/jenkins:lts-jdk17 +USER root +RUN jenkins-plugin-cli --plugins artifact-manager-s3 aws-credentials +USER jenkins +``` + +The `artifact-manager-s3` plugin brings in the AWS credentials support it needs; listing `aws-credentials` explicitly keeps the credential type available. + +Build and start Jenkins on the same Docker network as RustFS: + +```bash +docker build -t jenkins-rustfs . +docker run -d --name jenkins --network oo-rustfs_default \ + -p 8080:8080 -v jenkins-home:/var/jenkins_home jenkins-rustfs +``` + +Complete the setup wizard, then create an AWS credential: **Manage Jenkins → Credentials → global → Add Credentials**, kind **AWS Credential**, ID `rustfs-creds`, with your RustFS access key and secret key. + +## 2. Configure the artifact manager + +Open **Manage Jenkins → AWS Configuration** (from the `aws-global-configuration` plugin) and set: + +- **Region name**: `us-east-1` +- **Credentials**: `rustfs-creds` + +Open **Manage Jenkins → System** and locate the **Artifact Management for Builds** section. Select **Cloud Provider Amazon S3**, and fill in the S3 configuration: + +- **S3 Bucket Name**: `jenkins-artifacts` +- **S3 Bucket Region**: `us-east-1` +- **Base Prefix**: `artifacts/` +- **Custom Endpoint**: `:9000` (for example `rustfs:9000` inside the Compose network) +- **Custom Signing Region**: `us-east-1` +- **Use Path Style URL**: enabled +- **Use Insecure HTTP**: enabled +- **Disable Session Token**: enabled + +Path-style addressing and plain HTTP are required for a non-AWS endpoint without TLS. Disabling the session token stops the plugin from calling AWS STS, which a plain access key pair cannot answer. Click **Validate S3 Bucket configuration** to confirm the settings, then save. + +## 3. Run a job that archives an artifact + +Create a freestyle job (or pipeline) that produces a file and archives it: + +```groovy +pipeline { + agent any + stages { + stage('Build') { + steps { + sh 'echo "jenkins artifact stored on rustfs" > report.txt' + } + } + } + post { + always { + archiveArtifacts 'report.txt' + } + } +} +``` + +Run the build and wait for it to finish. The artifact upload goes to RustFS transparently — the job configuration does not mention S3 at all. + +## 4. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/jenkins-artifacts/ -r +``` + +The artifact is stored under the prefix, organized by job and build number: + +```text +artifacts/s3-artifacts-demo/3/artifacts/report.txt +``` + +![Jenkins artifacts stored in the RustFS Console](./images/rustfs-jenkins-artifacts.png) + +Downloading the artifact from the build page reads it back from RustFS. + +## 5. Stop or reset the deployment + +Stop Jenkins while keeping the data: + +```bash +docker rm -f jenkins +``` + +The artifacts stay in the `jenkins-artifacts` bucket. To delete them, remove the bucket: + +```bash +rc rb rustfs/jenkins-artifacts --force +``` + +## Troubleshooting + +### `StsException: The security token included in the request is invalid` + +The plugin is calling AWS STS to obtain session credentials. Enable **Disable Session Token** in the S3 configuration — a static access key pair cannot answer an STS call. + +### `UnknownHostException: jenkins-artifacts.rustfs` + +The plugin is using virtual-hosted addressing. Enable **Use Path Style URL** — RustFS resolves buckets from the URL path, not the hostname. + +### `No valid session credentials` or empty credential errors + +Confirm the AWS Configuration page has the credential selected and saved **before** the S3 bucket settings are used, and that the credential ID matches the one you created. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Jenkins integrations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Artifact Manager on S3 plugin documentation](https://plugins.jenkins.io/artifact-manager-s3/) for stash support and cleanup options. diff --git a/content/de/developer/integration/devops/meta.json b/content/de/developer/integration/devops/meta.json index 61ad9fb8..67c84c6e 100644 --- a/content/de/developer/integration/devops/meta.json +++ b/content/de/developer/integration/devops/meta.json @@ -3,6 +3,7 @@ "pages": [ "elasticsearch", "gitea", + "jenkins", "terraform" ] } diff --git a/content/de/developer/integration/index.md b/content/de/developer/integration/index.md index 868de8b2..e409e2e8 100644 --- a/content/de/developer/integration/index.md +++ b/content/de/developer/integration/index.md @@ -9,10 +9,10 @@ Use this section to connect **RustFS** to infrastructure and application platfor - [Reverse Proxy](./reverse-proxy/index.md) covers Nginx, Traefik, Caddy, and HAProxy. - [Backup](./backup/index.md) covers Restic and Longhorn. -- [Datenanalyse](./big-data/index.md) covers Iceberg. -- [Observability](./observability/index.md) covers OpenObserve. +- [Datenanalyse](./big-data/index.md) covers analytics systems including ClickHouse, Doris, Iceberg, Milvus, OpenDAL, and Zeppelin. +- [Observability](./observability/index.md) covers telemetry systems including Fluentd, OpenObserve, OpenTelemetry, Thanos, and Tempo. - [Others](./others/index.md) covers the community-driven capo SDK for Python. - [Registry](./registry/index.md) covers Harbor. -- [DevOps](./devops/index.md) covers Elasticsearch, Gitea, and Terraform. +- [DevOps](./devops/index.md) covers Elasticsearch, Gitea, Jenkins, and Terraform. Each guide identifies the RustFS endpoint and addressing requirements to use when configuring the integrating system. \ No newline at end of file diff --git a/content/de/developer/integration/observability/fluentd.md b/content/de/developer/integration/observability/fluentd.md new file mode 100644 index 00000000..4607c0ec --- /dev/null +++ b/content/de/developer/integration/observability/fluentd.md @@ -0,0 +1,152 @@ +--- +title: "Fluentd" +description: "Ship Fluentd log events to RustFS with the S3 output plugin." +--- + +This guide connects [Fluentd](https://github.com/fluent/fluentd) — the open-source data collector — to **RustFS** through the `out_s3` output plugin. You will run Fluentd with a tail source, buffer log events, and verify that the flushed objects are stored in RustFS. The workflow was verified with `fluent/fluentd:v1.17-1`, `fluent-plugin-s3` 1.8.6, and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + App["Application"] -->|"writes lines"| File["app.log"] + File -->|tail| Fluentd["Fluentd"] + Fluentd -->|"gzip objects"| RustFS["RustFS :9000"] +``` + +The tail source reads new lines from `app.log` and hands them to the S3 output, which buffers events on a time key and uploads a gzip object per flush window. + +## 1. Create the project files + +Create the bucket first — Fluentd does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/fluentd-data +``` + +The official Fluentd image runs as a non-root user and cannot install gems at startup, so build a small image with the S3 plugin: + +```dockerfile title="Dockerfile" +FROM fluent/fluentd:v1.17-1 +USER root +RUN gem install fluent-plugin-s3 --no-document +USER fluent +``` + +Create the Fluentd configuration, replacing both credential placeholders: + +```nginx title="fluent.conf" + + @type tail + path /var/log/app.log + pos_file /var/log/app.log.pos + tag rustfs.demo + + @type none + + + + + @type s3 + aws_key_id + aws_sec_key + s3_bucket fluentd-data + s3_endpoint http://rustfs:9000/ + s3_region us-east-1 + force_path_style true + path fluentd-logs + + @type memory + timekey 30s + timekey_wait 0s + flush_mode immediate + + +``` + +`force_path_style true` is required — without it the plugin constructs `fluentd-data.rustfs` as a hostname and every request fails with a DNS error. Inside the Compose network the hostname is `rustfs`; from the host use `http://localhost:9000/`. + +Build the image and start Fluentd on the same Docker network as RustFS: + +```bash +docker build -t fluentd-rustfs . +mkdir -p logs +docker run -d --name fluentd --network oo-rustfs_default \ + -v "$PWD/fluent.conf":/fluentd/etc/fluent.conf:ro \ + -v "$PWD/logs":/var/log fluentd-rustfs +``` + +## 2. Produce log events + +Append lines to the watched file: + +```bash +echo "rustfs fluentd demo line 1" >> logs/app.log +echo "rustfs fluentd demo line 2" >> logs/app.log +``` + +With a 30-second time key and immediate flush mode, each window uploads one gzip object shortly after it closes. Wait about a minute. + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/fluentd-data/ -r +``` + +Each flush window produces one gzipped object: + +```text +fluentd-logs20260921154400_0.gz +fluentd-logs20260921154400_1.gz +``` + +Read one object back to confirm the events are intact: + +```bash +rc cat rustfs/fluentd-data/fluentd-logs20260921154400_0.gz | gunzip +``` + +![Fluentd log objects stored in the RustFS Console](./images/rustfs-fluentd-logs.png) + +## 4. Stop or reset the deployment + +Stop Fluentd while keeping the data: + +```bash +docker rm -f fluentd +``` + +The objects stay in the `fluentd-data` bucket. To delete them, remove the bucket: + +```bash +rc rb rustfs/fluentd-data --force +``` + +## Troubleshooting + +### `Unknown output plugin 's3'` + +The plugin is not installed. Confirm the Dockerfile installs `fluent-plugin-s3` as `root` before switching back to the `fluent` user — installing at container start as the default user fails with permission errors. + +### `Failed to open TCP connection to fluentd-data.rustfs` + +Virtual-hosted addressing is in use. Add `force_path_style true` to the `s3` output so the bucket stays in the URL path. + +### The worker crashes in a restart loop + +The output fails hard when the bucket does not exist. Create `fluentd-data` before starting Fluentd and check the startup log: + +```bash +docker logs fluentd | grep -iE "error|bucket" | tail +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Fluentd outputs. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [fluent-plugin-s3 documentation](https://github.com/fluent/fluent-plugin-s3) for object key formats and compression options. diff --git a/content/de/developer/integration/observability/images/rustfs-fluentd-logs.png b/content/de/developer/integration/observability/images/rustfs-fluentd-logs.png new file mode 100644 index 00000000..3ab768ca Binary files /dev/null and b/content/de/developer/integration/observability/images/rustfs-fluentd-logs.png differ diff --git a/content/de/developer/integration/observability/images/rustfs-otel-logs.png b/content/de/developer/integration/observability/images/rustfs-otel-logs.png new file mode 100644 index 00000000..2f68a22a Binary files /dev/null and b/content/de/developer/integration/observability/images/rustfs-otel-logs.png differ diff --git a/content/de/developer/integration/observability/index.md b/content/de/developer/integration/observability/index.md index 3a1a4151..99a78611 100644 --- a/content/de/developer/integration/observability/index.md +++ b/content/de/developer/integration/observability/index.md @@ -7,7 +7,9 @@ Nutzen Sie **RustFS** als Objektspeicher-Layer für Observability-Plattformen, d ## Plattformen +- [Fluentd](./fluentd.md) - [OpenObserve](./openobserve.md) +- [OpenTelemetry](./opentelemetry.md) - [Loki](./loki.md) - [Tempo](./tempo.md) - [Thanos](./thanos.md) diff --git a/content/de/developer/integration/observability/meta.json b/content/de/developer/integration/observability/meta.json index 03f10c67..d6ddebd7 100644 --- a/content/de/developer/integration/observability/meta.json +++ b/content/de/developer/integration/observability/meta.json @@ -1,8 +1,10 @@ { "title": "Observability", "pages": [ - "openobserve", + "fluentd", "loki", + "openobserve", + "opentelemetry", "tempo", "thanos" ] diff --git a/content/de/developer/integration/observability/opentelemetry.md b/content/de/developer/integration/observability/opentelemetry.md new file mode 100644 index 00000000..283df413 --- /dev/null +++ b/content/de/developer/integration/observability/opentelemetry.md @@ -0,0 +1,145 @@ +--- +title: "OpenTelemetry" +description: "Export OpenTelemetry Collector logs to RustFS with the AWS S3 exporter." +--- + +This guide connects the [OpenTelemetry Collector](https://opentelemetry.io/docs/collector/) — the CNCF telemetry pipeline — to **RustFS** through the collector's `awss3` exporter. You will run the contrib collector with a `filelog` receiver, ship the log lines of a file into the `otel-data` bucket, and verify the partitioned objects in RustFS. The workflow was verified with `otel/opentelemetry-collector-contrib:0.138.0` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Log["Application log file"] -->|filelog| Collector["OTel Collector"] + Collector -->|awss3| RustFS["RustFS :9000"] +``` + +The `filelog` receiver tails the log file and the `awss3` exporter uploads batches to the bucket, partitioned into time-based prefixes. Metrics and traces can be routed through the same exporter with their own pipelines. + +## 1. Create the project files + +Create the bucket first — the exporter does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/otel-data +``` + +Create the collector configuration, replacing both credential placeholders: + +```yaml title="config.yaml" +receivers: + filelog: + include: [/var/log/app.log] + start_at: beginning + +exporters: + awss3: + s3uploader: + region: us-east-1 + s3_bucket: otel-data + endpoint: http://rustfs:9000 + s3_force_path_style: true + disable_ssl: true + file_prefix: logs/app + marshaler: body + +service: + pipelines: + logs: + receivers: [filelog] + exporters: [awss3] +``` + +The exporter reads credentials from the standard `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, and `AWS_REGION` environment variables. `marshaler: body` writes each log line as plain text; omit it to store OTLP JSON instead. `s3_force_path_style` and `disable_ssl` are required for a plain-HTTP, non-AWS endpoint. + +Create the Compose file: + +```yaml title="compose.yaml" +services: + collector: + image: otel/opentelemetry-collector-contrib:0.138.0 + command: ["--config=/etc/otelcol-contrib/config.yaml"] + environment: + AWS_ACCESS_KEY_ID: + AWS_SECRET_ACCESS_KEY: + AWS_REGION: us-east-1 + volumes: + - ./config.yaml:/etc/otelcol-contrib/config.yaml:ro + - ./app.log:/var/log/app.log + networks: + - otel + +networks: + otel: +``` + +## 2. Start the collector and produce logs + +Start the stack and append lines to the watched file: + +```bash +echo "otel demo log line one" > app.log +docker compose up -d +echo "second line after start" >> app.log +``` + +The collector tails the file from the beginning and uploads each buffer when the partition rolls over, so allow about a minute after the last line before checking. + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/otel-data/ -r +``` + +Log records are uploaded under time-based partitions with the configured file prefix: + +```text +year=2026/month=09/day=21/hour=15/minute=53/logs/applogs_288361608.txt +``` + +Read one object back to confirm the lines arrived intact: + +```bash +rc cat rustfs/otel-data/year=2026/month=09/day=21/hour=15/minute=53/logs/applogs_288361608.txt +``` + +```text +otel demo log line one +second line after start +``` + +![OpenTelemetry log objects stored in the RustFS Console](./images/rustfs-otel-logs.png) + +## 4. Stop or reset the deployment + +Stop the collector while keeping the data: + +```bash +docker compose down +``` + +The objects stay in the `otel-data` bucket. To delete them, remove the bucket: + +```bash +rc rb rustfs/otel-data --force +``` + +## Troubleshooting + +### `has invalid keys` when the collector starts + +The `awss3` exporter schema differs between collector releases. Version 0.138 nests the upload settings under `s3uploader` as shown above; the endpoint key is `endpoint` (not `s3_endpoint`). Newer releases move these keys to the top level — check the README for your exact collector version. + +### Nothing appears in the bucket + +Confirm the credentials environment variables are set on the collector container, that `s3_force_path_style` is `true`, and that the exporter can reach `http://rustfs:9000` from inside the Compose network. Enable the `debug` exporter on the same pipeline to see whether records flow at all. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional collector operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [AWS S3 exporter documentation](https://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/exporter/awss3exporter) to add metrics and traces pipelines. diff --git a/content/en/developer/integration/big-data/clickhouse.md b/content/en/developer/integration/big-data/clickhouse.md new file mode 100644 index 00000000..ac9d8d4b --- /dev/null +++ b/content/en/developer/integration/big-data/clickhouse.md @@ -0,0 +1,160 @@ +--- +title: "ClickHouse" +description: "Run ClickHouse with an S3 disk backed by RustFS for MergeTree table data." +--- + +This guide connects [ClickHouse](https://github.com/ClickHouse/ClickHouse) — the real-time OLAP database — to **RustFS** through ClickHouse's S3 disk storage policy. You will start a ClickHouse server with Docker, create a MergeTree table that stores its parts on RustFS, insert rows, and verify that the table data lives in the bucket. The workflow was verified with `clickhouse/clickhouse-server:25.8` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["SQL client"] -->|"queries"| CH["ClickHouse :8123"] + CH -->|"MergeTree parts"| RustFS["RustFS :9000"] +``` + +The `rustfs` disk is a ClickHouse S3 disk pointed at the `clickhouse-data` bucket. Tables created with the matching storage policy write their parts — data, index, and checksum files — to the bucket instead of the local filesystem. + +## 1. Create the project files + +Create the bucket first — ClickHouse does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/clickhouse-data +``` + +Create the storage configuration, replacing both credential placeholders: + +```xml title="storage.xml" + + + + + s3 + http://rustfs:9000/clickhouse-data/ + + + + + + + +
+ rustfs +
+
+
+
+
+
+``` + +The endpoint must end with `/` and includes the bucket name as the first path segment. Inside the Compose network the hostname is `rustfs`; from the host use `http://localhost:9000/clickhouse-data/`. + +Start ClickHouse with the configuration mounted: + +```bash +docker run -d --name clickhouse --network oo-rustfs_default \ + -p 8123:8123 \ + -e CLICKHOUSE_PASSWORD= \ + -v "$PWD/storage.xml":/etc/clickhouse-server/config.d/storage.xml:ro \ + clickhouse/clickhouse-server:25.8 +``` + +## 2. Create a table on the S3 disk + +Wait for the HTTP interface, then create a database and a MergeTree table with the storage policy: + +```bash +curl "http://localhost:8123/?password=" \ + --data-binary "CREATE DATABASE rustfs_demo" + +curl "http://localhost:8123/?password=" \ + --data-binary "CREATE TABLE rustfs_demo.events + (id UInt32, name String) + ENGINE = MergeTree ORDER BY id + SETTINGS storage_policy = 'rustfs_policy'" + +curl "http://localhost:8123/?password=" \ + --data-binary "INSERT INTO rustfs_demo.events + VALUES (1, 'clickhouse-on-rustfs'), (2, 'second')" +``` + +Read the rows back and confirm ClickHouse reports the part on the `rustfs` disk: + +```bash +curl "http://localhost:8123/?password=" \ + --data-binary "SELECT count(), any(name) FROM rustfs_demo.events" + +curl "http://localhost:8123/?password=" \ + --data-binary "SELECT name, disk_name FROM system.parts + WHERE database = 'rustfs_demo' AND active" +``` + +```text +2 clickhouse-on-rustfs +all_1_1_0 rustfs +``` + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/clickhouse-data/ -r +``` + +ClickHouse writes each part as content-addressed blobs. The output contains several small objects, and the count grows as more parts are written: + +```text +dtg/hpsyncexixvdnsgorvseobogcgowg +dzp/zfblobhsatzdveqdrsfqcupkdehja +izg/gvhchqobrizpkdftuvlmakkoutwps +``` + +![ClickHouse parts stored in the RustFS Console](./images/rustfs-clickhouse-disk.png) + +The data survives a container restart because the parts live in RustFS: + +```bash +docker restart clickhouse +curl "http://localhost:8123/?password=" \ + --data-binary "SELECT count() FROM rustfs_demo.events" +``` + +## 4. Stop or reset the deployment + +Stop the server while keeping the data: + +```bash +docker rm -f clickhouse +``` + +The parts stay in the `clickhouse-data` bucket and the table can be queried again after the next start. To delete the data, remove the bucket: + +```bash +rc rb rustfs/clickhouse-data --force +``` + +## Troubleshooting + +### `REQUIRED_PASSWORD` on every query + +ClickHouse 25.8 images require a password for the `default` user. Set `CLICKHOUSE_PASSWORD` on the container and pass the same value as the `password` query parameter, as shown above. + +### Table creation fails with a disk or endpoint error + +Confirm that the bucket exists before the table is created, that the endpoint ends with `/`, and that the credentials match the RustFS deployment. Check the server log for the underlying S3 error: + +```bash +docker logs clickhouse | grep -i s3 | tail +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional ClickHouse operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [ClickHouse S3 disk documentation](https://clickhouse.com/docs/engines/table-engines/mergetree-family/mergetree#table_engine-mergetree-s3) to add a cache disk or a tiered hot/cold policy. diff --git a/content/en/developer/integration/big-data/doris.md b/content/en/developer/integration/big-data/doris.md new file mode 100644 index 00000000..1c3665b5 --- /dev/null +++ b/content/en/developer/integration/big-data/doris.md @@ -0,0 +1,161 @@ +--- +title: "Apache Doris" +description: "Back up Apache Doris tables to RustFS through an S3 repository and restore them." +--- + +This guide connects [Apache Doris](https://github.com/apache/doris) — the real-time analytical data warehouse — to **RustFS** through an S3 backup repository. You will start an all-in-one Doris container, create an S3 repository pointing at a RustFS bucket, back up a table, drop it, and restore it from RustFS. The workflow was verified with `apache/doris:all-in-one-4.1.3` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["SQL client"] -->|"queries"| Doris["Doris FE/BE"] + Doris -->|"BACKUP / RESTORE"| RustFS["RustFS :9000"] +``` + +The repository is a named S3 location under the `doris-backups` bucket. `BACKUP SNAPSHOT` uploads table metadata and tablet data files; `RESTORE SNAPSHOT` downloads them into a new table. + +## 1. Start Doris and create the repository + +Create the bucket first — Doris does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/doris-backups +``` + +Start the all-in-one container on the same Docker network as RustFS: + +```bash +docker run -d --name doris --network oo-rustfs_default \ + -p 8030:8030 -p 9030:9030 apache/doris:all-in-one-4.1.3 +``` + +Wait for the frontend to become healthy, then connect with the MySQL protocol (port 9030, user `root`, no password in the all-in-one image). + +Create a test table with rows: + +```sql +CREATE DATABASE rustfs_demo; +CREATE TABLE rustfs_demo.events + (id INT, name VARCHAR(50)) + DISTRIBUTED BY HASH(id) BUCKETS 1 + PROPERTIES ("replication_num" = "1"); +INSERT INTO rustfs_demo.events VALUES (1, 'doris-on-rustfs'), (2, 'backup-test'); +``` + +Create the S3 repository, replacing the endpoint with the IP address of the RustFS container and both credential placeholders: + +```sql +CREATE REPOSITORY `rustfs_repo` + WITH S3 + ON LOCATION "s3://doris-backups/rustfs-repo" + PROPERTIES ( + "AWS_ENDPOINT" = "http://:9000", + "AWS_ACCESS_KEY" = "", + "AWS_SECRET_KEY" = "", + "AWS_REGION" = "us-east-1", + "AWS_PATH_STYLE_ACCESS" = "true" + ); +``` + +Doris 4.1 resolves the bucket into the endpoint hostname even with `AWS_PATH_STYLE_ACCESS` enabled, so a hostname endpoint fails with `UnknownHostException: doris-backups.rustfs`. Using the container IP address forces path-style requests and works; `SHOW REPOSITORIES` confirms the repository registered with an empty `ErrMsg`. + +## 2. Back up a table to RustFS + +Take a snapshot of the table: + +```sql +BACKUP SNAPSHOT rustfs_demo.demo_snapshot + TO rustfs_repo + ON (events); +``` + +The statement returns immediately; the backup job runs in the background. Watch its state: + +```sql +SHOW BACKUP; +``` + +Wait until `State` reaches `FINISHED` — the snapshot metadata and tablet data files are now objects in the bucket. + +## 3. Verify the backup in RustFS + +List the bucket: + +```bash +rc ls rustfs/doris-backups/ -r +``` + +The repository stores a repository descriptor, the snapshot metadata, and the tablet files: + +```text +rustfs-repo/__palo_repository_rustfs_repo/__repo_info +rustfs-repo/__palo_repository_rustfs_repo/__ss_demo_snapshot/__meta.d50ecf9b... +rustfs-repo/__palo_repository_rustfs_repo/__ss_demo_snapshot/__ss_content/.../...dat... +``` + +![Doris backup objects stored in the RustFS Console](./images/rustfs-doris-backup.png) + +## 4. Restore the table from RustFS + +Drop the table and restore it from the snapshot. The timestamp comes from the snapshot name shown by `SHOW SNAPSHOT ON REPOSITORY rustfs_repo;`: + +```sql +DROP TABLE rustfs_demo.events; + +RESTORE SNAPSHOT rustfs_demo.demo_snapshot + FROM rustfs_repo + ON (events) + PROPERTIES ( + "backup_timestamp" = "2026-09-21-16-35-12", + "replication_num" = "1" + ); +``` + +Wait for the restore job to finish and confirm the data: + +```sql +SHOW RESTORE; +SELECT count(*) FROM rustfs_demo.events; +``` + +```text +2 +``` + +## 5. Stop or reset the deployment + +Stop Doris while keeping the data: + +```bash +docker rm -f doris +``` + +The backup stays in the `doris-backups` bucket and can be restored into any Doris cluster that registers the same repository. To delete it, remove the bucket: + +```bash +rc rb rustfs/doris-backups --force +``` + +## Troubleshooting + +### `UnknownHostException: doris-backups.rustfs` when creating the repository + +Doris is building a virtual-hosted hostname from the bucket and endpoint. Use the RustFS container IP address in `AWS_ENDPOINT` together with `AWS_PATH_STYLE_ACCESS = "true"`, as shown above. + +### The backup stays in `SNAPSHOTING` for a long time + +The backend uploads the tablet files. Confirm the backend is healthy (`SHOW BACKENDS;`) and can reach the endpoint; the all-in-one image needs a minute or two after start before both processes report ready. + +### `Failed to create repository: ... file status` + +The bucket does not exist or the credentials are wrong. Create `doris-backups` with `rc mb` and re-check the access key pair. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Doris operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Doris backup and restore documentation](https://doris.apache.org/docs/data-operate/backup-restore/) to schedule periodic snapshots. diff --git a/content/en/developer/integration/big-data/images/rustfs-clickhouse-disk.png b/content/en/developer/integration/big-data/images/rustfs-clickhouse-disk.png new file mode 100644 index 00000000..9225bf66 Binary files /dev/null and b/content/en/developer/integration/big-data/images/rustfs-clickhouse-disk.png differ diff --git a/content/en/developer/integration/big-data/images/rustfs-doris-backup.png b/content/en/developer/integration/big-data/images/rustfs-doris-backup.png new file mode 100644 index 00000000..9a936276 Binary files /dev/null and b/content/en/developer/integration/big-data/images/rustfs-doris-backup.png differ diff --git a/content/en/developer/integration/big-data/images/rustfs-opendal-objects.png b/content/en/developer/integration/big-data/images/rustfs-opendal-objects.png new file mode 100644 index 00000000..d7129423 Binary files /dev/null and b/content/en/developer/integration/big-data/images/rustfs-opendal-objects.png differ diff --git a/content/en/developer/integration/big-data/images/rustfs-zeppelin-notebook.png b/content/en/developer/integration/big-data/images/rustfs-zeppelin-notebook.png new file mode 100644 index 00000000..9c5fe782 Binary files /dev/null and b/content/en/developer/integration/big-data/images/rustfs-zeppelin-notebook.png differ diff --git a/content/en/developer/integration/big-data/index.md b/content/en/developer/integration/big-data/index.md index f207d9b7..bb96ab87 100644 --- a/content/en/developer/integration/big-data/index.md +++ b/content/en/developer/integration/big-data/index.md @@ -7,11 +7,14 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo ## Systems +- [ClickHouse](./clickhouse.md) - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) - [Milvus](./milvus.md) - [MLflow](./mlflow.md) +- [OpenDAL](./opendal.md) - [DuckDB](./duckdb.md) +- [Doris](./doris.md) - [InfluxDB](./influxdb.md) - [Spark](./spark.md) - [Flink](./flink.md) diff --git a/content/en/developer/integration/big-data/meta.json b/content/en/developer/integration/big-data/meta.json index 87e29f91..ee7d69b5 100644 --- a/content/en/developer/integration/big-data/meta.json +++ b/content/en/developer/integration/big-data/meta.json @@ -1,14 +1,18 @@ { "title": "Data Analytics", "pages": [ + "clickhouse", "iceberg", "pyiceberg", "milvus", "mlflow", + "opendal", "duckdb", + "doris", "influxdb", "spark", "flink", - "trino" + "trino", + "zeppelin" ] } diff --git a/content/en/developer/integration/big-data/opendal.md b/content/en/developer/integration/big-data/opendal.md new file mode 100644 index 00000000..074b7475 --- /dev/null +++ b/content/en/developer/integration/big-data/opendal.md @@ -0,0 +1,108 @@ +--- +title: "OpenDAL" +description: "Access RustFS objects from applications through the Apache OpenDAL data access layer." +--- + +This guide connects [Apache OpenDAL](https://github.com/apache/opendal) — the unified data access layer — to **RustFS** through its `s3` service. You will run the OpenDAL Python binding against RustFS, write and read an object, list a prefix, and delete it. The workflow was verified with the `opendal` Python package 0.46 on `python:3.12-slim` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and Python 3.9 or later. OpenDAL supports the same `s3` service from Rust, Java, Node.js, and Go bindings with equivalent settings. + +## Architecture + +```mermaid +flowchart LR + App["Application"] -->|"Operator API"| OpenDAL["OpenDAL"] + OpenDAL -->|"s3 service"| RustFS["RustFS :9000"] +``` + +An OpenDAL `Operator` bound to the `s3` service exposes one uniform API — `write`, `read`, `stat`, `list`, `delete` — over the bucket, so the same code runs against S3, RustFS, or any other supported service by changing the connection settings. + +## 1. Set up the project + +Install the Python binding: + +```bash +pip install opendal +``` + +Create the script, replacing all connection placeholders: + +```python title="opendal_demo.py" +import opendal + +op = opendal.Operator( + "s3", + endpoint="http://:9000", + bucket="my-bucket", + access_key_id="", + secret_access_key="", + region="us-east-1", +) +op.write("opendal-demo/hello.txt", b"hello from opendal against rustfs") +print("read-back:", op.read("opendal-demo/hello.txt")) +print("content_length:", op.stat("opendal-demo/hello.txt").content_length) +for entry in op.list("opendal-demo/"): + print("listed:", entry.path) +op.delete("opendal-demo/hello.txt") +print("deleted:", not op.exists("opendal-demo/hello.txt")) +``` + +The credential option names are `access_key_id` and `secret_access_key` — the shorter `access_key` names do not exist and fail with a signing error. Path-style addressing is the default for non-AWS endpoints. + +## 2. Run the demo + +Run the script from a machine that can reach RustFS: + +```bash +python opendal_demo.py +``` + +```text +read-back: b"hello from opendal against rustfs" +content_length: 33 +listed: opendal-demo/hello.txt +deleted: True +``` + +The round trip exercises the full object lifecycle: write uploads the bytes, `read` fetches them back, `stat` returns the object size, `list` enumerates the prefix, and `delete` removes the object. + +## 3. Verify objects in RustFS + +Comment out the final `op.delete` line, run the script again, and list the prefix in RustFS: + +```bash +rc ls rustfs/my-bucket/opendal-demo/ -r +``` + +```text +hello.txt +data/rows.csv +``` + +![OpenDAL objects stored in the RustFS Console](./images/rustfs-opendal-objects.png) + +The objects visible in the RustFS Console are exactly the paths the OpenDAL API wrote. + +## 4. Stop or reset + +OpenDAL is a library and holds no state of its own. To clean up the demo objects: + +```bash +rc rm rustfs/my-bucket/opendal-demo/ --recursive --force +``` + +## Troubleshooting + +### `failed to load signing credential` + +The operator received no usable credentials. Use the exact option names `access_key_id` and `secret_access_key`; other spellings are silently ignored and signing then fails. + +### Connection or DNS errors on write + +Confirm the endpoint includes the scheme and port and is reachable from the application. Inside a Compose network the hostname is `rustfs`; from the host use `http://localhost:9000`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional OpenDAL operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [OpenDAL documentation](https://opendal.apache.org/docs/) to use the same operator from Rust, Java, or Node.js. diff --git a/content/en/developer/integration/big-data/zeppelin.md b/content/en/developer/integration/big-data/zeppelin.md new file mode 100644 index 00000000..e7f47375 --- /dev/null +++ b/content/en/developer/integration/big-data/zeppelin.md @@ -0,0 +1,109 @@ +--- +title: "Apache Zeppelin" +description: "Store Apache Zeppelin notebooks in RustFS through the S3 notebook repository." +--- + +This guide connects [Apache Zeppelin](https://github.com/apache/zeppelin) — the web-based notebook for data analytics — to **RustFS** through Zeppelin's S3 notebook storage. You will start Zeppelin pointed at a RustFS bucket, create a note, and verify that the notebook file is stored in RustFS and survives a restart. The workflow was verified with `apache/zeppelin:0.12.0` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Browser["Browser"] -->|"notebook edits"| Z["Zeppelin :8080"] + Z -->|".zpln files"| RustFS["RustFS :9000"] +``` + +With the `S3NotebookRepo` storage class, every note is persisted as a `.zpln` JSON file under the `user/notebook/` prefix of the bucket. Zeppelin reads and writes the bucket directly, so notes survive container restarts and can be shared between instances. + +## 1. Start Zeppelin + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +AWS_ACCESS_KEY_ID= +AWS_SECRET_ACCESS_KEY= +``` + +Start Zeppelin on the same Docker network as RustFS with the S3 storage settings: + +```bash +docker run -d --name zeppelin --network oo-rustfs_default \ + -p 8080:8080 \ + -e AWS_ACCESS_KEY_ID \ + -e AWS_SECRET_ACCESS_KEY \ + -e ZEPPELIN_NOTEBOOK_STORAGE=org.apache.zeppelin.notebook.repo.S3NotebookRepo \ + -e ZEPPELIN_NOTEBOOK_S3_BUCKET=my-bucket \ + -e ZEPPELIN_NOTEBOOK_S3_ENDPOINT=http://rustfs:9000 \ + -e ZEPPELIN_NOTEBOOK_S3_PATH_STYLE_ACCESS=true \ + apache/zeppelin:0.12.0 +``` + +Zeppelin reads `ZEPPELIN_*` environment variables as configuration properties, so no `zeppelin-site.xml` edit is needed. This guide uses the existing `my-bucket`; notes land under its `user/notebook/` prefix, which the S3 storage creates on demand. + +## 2. Create a note + +Wait for the UI on `http://localhost:8080`, then create a note named `rustfs-demo` in the notebook list and add a paragraph, or use the REST API: + +```bash +NOTE=$(curl -s -X POST "http://localhost:8080/api/notebook" \ + -H "Content-Type: application/json" \ + -d '{"name": "rustfs-demo"}' | python3 -c "import json,sys; print(json.load(sys.stdin)['body'])") +echo "note id: $NOTE" +``` + +## 3. Verify the notebook in RustFS + +List the notebook prefix: + +```bash +rc ls rustfs/my-bucket/user/notebook/ -r +``` + +The note is stored as a JSON file named after the note and its ID: + +```text +user/notebook/rustfs-demo_2N4PY7UY5.zpln +``` + +![Zeppelin notebooks stored in the RustFS Console](./images/rustfs-zeppelin-notebook.png) + +Notes survive a restart because Zeppelin loads them from the bucket: + +```bash +docker restart zeppelin +curl -s "http://localhost:8080/api/notebook" | head -c 200 +``` + +The note list contains `2N4PY7UY5` again after the restart. + +## 4. Stop or reset the deployment + +Stop Zeppelin while keeping the notes: + +```bash +docker rm -f zeppelin +``` + +The notes stay in the `user/notebook/` prefix of `my-bucket`. To delete them, remove the prefix: + +```bash +rc rm rustfs/my-bucket/user/notebook/ --recursive --force +``` + +## Troubleshooting + +### Zeppelin starts but notes never appear in the bucket + +Confirm the three `ZEPPELIN_NOTEBOOK_S3_*` variables are set and that the credentials environment variables reach the container — the S3 repository is initialized at startup, so the container must be recreated after any change. + +### `UnknownHostException: my-bucket.rustfs` + +The path-style flag was not picked up. Keep `ZEPPELIN_NOTEBOOK_S3_PATH_STYLE_ACCESS=true` exactly as shown; with path-style disabled Zeppelin treats the bucket as a hostname. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Zeppelin operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Zeppelin storage documentation](https://zeppelin.apache.org/docs/latest/setup/storage/storage.html#notebook-storage-in-s3) to organize notes in per-user prefixes. diff --git a/content/en/developer/integration/devops/images/rustfs-jenkins-artifacts.png b/content/en/developer/integration/devops/images/rustfs-jenkins-artifacts.png new file mode 100644 index 00000000..19c8cf87 Binary files /dev/null and b/content/en/developer/integration/devops/images/rustfs-jenkins-artifacts.png differ diff --git a/content/en/developer/integration/devops/index.md b/content/en/developer/integration/devops/index.md index 95f6261e..ad31918d 100644 --- a/content/en/developer/integration/devops/index.md +++ b/content/en/developer/integration/devops/index.md @@ -9,6 +9,7 @@ Use **RustFS** as the object storage layer for DevOps platforms and infrastructu - [Elasticsearch](./elasticsearch.md) - [Gitea](./gitea.md) +- [Jenkins](./jenkins.md) - [Terraform](./terraform.md) Keep artifacts, state, and telemetry in dedicated buckets, and use credentials scoped to the required bucket operations. diff --git a/content/en/developer/integration/devops/jenkins.md b/content/en/developer/integration/devops/jenkins.md new file mode 100644 index 00000000..1d78f31b --- /dev/null +++ b/content/en/developer/integration/devops/jenkins.md @@ -0,0 +1,144 @@ +--- +title: "Jenkins" +description: "Store Jenkins build artifacts in RustFS with the Artifact Manager on S3 plugin." +--- + +This guide connects [Jenkins](https://github.com/jenkinsci/jenkins) — the automation server — to **RustFS** through the Artifact Manager on S3 plugin. You will start Jenkins with the plugin, point its artifact manager at a RustFS bucket, run a job that archives an artifact, and verify that the artifact is stored in RustFS. The workflow was verified with `jenkins/jenkins:lts-jdk17` (Jenkins 2.5xx LTS) and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Dev["Developer"] -->|"trigger build"| J["Jenkins :8080"] + J -->|"archiveArtifacts"| RustFS["RustFS :9000"] +``` + +With the plugin active, every job that publishes artifacts through the standard `archiveArtifacts` step (or `stash`/`unstash`) uploads them to the `jenkins-artifacts` bucket under the configured prefix instead of storing them on the controller disk. + +## 1. Create the project files + +Create the bucket first — the plugin validates the bucket but does not create it: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/jenkins-artifacts +``` + +Create a Dockerfile that installs the plugin into the Jenkins LTS image: + +```dockerfile title="Dockerfile" +FROM jenkins/jenkins:lts-jdk17 +USER root +RUN jenkins-plugin-cli --plugins artifact-manager-s3 aws-credentials +USER jenkins +``` + +The `artifact-manager-s3` plugin brings in the AWS credentials support it needs; listing `aws-credentials` explicitly keeps the credential type available. + +Build and start Jenkins on the same Docker network as RustFS: + +```bash +docker build -t jenkins-rustfs . +docker run -d --name jenkins --network oo-rustfs_default \ + -p 8080:8080 -v jenkins-home:/var/jenkins_home jenkins-rustfs +``` + +Complete the setup wizard, then create an AWS credential: **Manage Jenkins → Credentials → global → Add Credentials**, kind **AWS Credential**, ID `rustfs-creds`, with your RustFS access key and secret key. + +## 2. Configure the artifact manager + +Open **Manage Jenkins → AWS Configuration** (from the `aws-global-configuration` plugin) and set: + +- **Region name**: `us-east-1` +- **Credentials**: `rustfs-creds` + +Open **Manage Jenkins → System** and locate the **Artifact Management for Builds** section. Select **Cloud Provider Amazon S3**, and fill in the S3 configuration: + +- **S3 Bucket Name**: `jenkins-artifacts` +- **S3 Bucket Region**: `us-east-1` +- **Base Prefix**: `artifacts/` +- **Custom Endpoint**: `:9000` (for example `rustfs:9000` inside the Compose network) +- **Custom Signing Region**: `us-east-1` +- **Use Path Style URL**: enabled +- **Use Insecure HTTP**: enabled +- **Disable Session Token**: enabled + +Path-style addressing and plain HTTP are required for a non-AWS endpoint without TLS. Disabling the session token stops the plugin from calling AWS STS, which a plain access key pair cannot answer. Click **Validate S3 Bucket configuration** to confirm the settings, then save. + +## 3. Run a job that archives an artifact + +Create a freestyle job (or pipeline) that produces a file and archives it: + +```groovy +pipeline { + agent any + stages { + stage('Build') { + steps { + sh 'echo "jenkins artifact stored on rustfs" > report.txt' + } + } + } + post { + always { + archiveArtifacts 'report.txt' + } + } +} +``` + +Run the build and wait for it to finish. The artifact upload goes to RustFS transparently — the job configuration does not mention S3 at all. + +## 4. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/jenkins-artifacts/ -r +``` + +The artifact is stored under the prefix, organized by job and build number: + +```text +artifacts/s3-artifacts-demo/3/artifacts/report.txt +``` + +![Jenkins artifacts stored in the RustFS Console](./images/rustfs-jenkins-artifacts.png) + +Downloading the artifact from the build page reads it back from RustFS. + +## 5. Stop or reset the deployment + +Stop Jenkins while keeping the data: + +```bash +docker rm -f jenkins +``` + +The artifacts stay in the `jenkins-artifacts` bucket. To delete them, remove the bucket: + +```bash +rc rb rustfs/jenkins-artifacts --force +``` + +## Troubleshooting + +### `StsException: The security token included in the request is invalid` + +The plugin is calling AWS STS to obtain session credentials. Enable **Disable Session Token** in the S3 configuration — a static access key pair cannot answer an STS call. + +### `UnknownHostException: jenkins-artifacts.rustfs` + +The plugin is using virtual-hosted addressing. Enable **Use Path Style URL** — RustFS resolves buckets from the URL path, not the hostname. + +### `No valid session credentials` or empty credential errors + +Confirm the AWS Configuration page has the credential selected and saved **before** the S3 bucket settings are used, and that the credential ID matches the one you created. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Jenkins integrations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Artifact Manager on S3 plugin documentation](https://plugins.jenkins.io/artifact-manager-s3/) for stash support and cleanup options. diff --git a/content/en/developer/integration/devops/meta.json b/content/en/developer/integration/devops/meta.json index 61ad9fb8..67c84c6e 100644 --- a/content/en/developer/integration/devops/meta.json +++ b/content/en/developer/integration/devops/meta.json @@ -3,6 +3,7 @@ "pages": [ "elasticsearch", "gitea", + "jenkins", "terraform" ] } diff --git a/content/en/developer/integration/index.md b/content/en/developer/integration/index.md index 5e604df9..38f2ac3d 100644 --- a/content/en/developer/integration/index.md +++ b/content/en/developer/integration/index.md @@ -9,10 +9,10 @@ Use this section to connect **RustFS** to infrastructure and application platfor - [Reverse Proxy](./reverse-proxy/index.md) covers Nginx, Traefik, Caddy, and HAProxy. - [Backup](./backup/index.md) covers Restic and Longhorn. -- [Data Analytics](./big-data/index.md) covers Iceberg. -- [Observability](./observability/index.md) covers OpenObserve. +- [Data Analytics](./big-data/index.md) covers analytics systems including ClickHouse, Doris, Iceberg, Milvus, OpenDAL, and Zeppelin. +- [Observability](./observability/index.md) covers telemetry systems including Fluentd, OpenObserve, OpenTelemetry, Thanos, and Tempo. - [Others](./others/index.md) covers the community-driven capo SDK for Python. - [Registry](./registry/index.md) covers Harbor. -- [DevOps](./devops/index.md) covers Elasticsearch, Gitea, and Terraform. +- [DevOps](./devops/index.md) covers Elasticsearch, Gitea, Jenkins, and Terraform. Each guide identifies the RustFS endpoint and addressing requirements to use when configuring the integrating system. \ No newline at end of file diff --git a/content/en/developer/integration/observability/fluentd.md b/content/en/developer/integration/observability/fluentd.md new file mode 100644 index 00000000..4607c0ec --- /dev/null +++ b/content/en/developer/integration/observability/fluentd.md @@ -0,0 +1,152 @@ +--- +title: "Fluentd" +description: "Ship Fluentd log events to RustFS with the S3 output plugin." +--- + +This guide connects [Fluentd](https://github.com/fluent/fluentd) — the open-source data collector — to **RustFS** through the `out_s3` output plugin. You will run Fluentd with a tail source, buffer log events, and verify that the flushed objects are stored in RustFS. The workflow was verified with `fluent/fluentd:v1.17-1`, `fluent-plugin-s3` 1.8.6, and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + App["Application"] -->|"writes lines"| File["app.log"] + File -->|tail| Fluentd["Fluentd"] + Fluentd -->|"gzip objects"| RustFS["RustFS :9000"] +``` + +The tail source reads new lines from `app.log` and hands them to the S3 output, which buffers events on a time key and uploads a gzip object per flush window. + +## 1. Create the project files + +Create the bucket first — Fluentd does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/fluentd-data +``` + +The official Fluentd image runs as a non-root user and cannot install gems at startup, so build a small image with the S3 plugin: + +```dockerfile title="Dockerfile" +FROM fluent/fluentd:v1.17-1 +USER root +RUN gem install fluent-plugin-s3 --no-document +USER fluent +``` + +Create the Fluentd configuration, replacing both credential placeholders: + +```nginx title="fluent.conf" + + @type tail + path /var/log/app.log + pos_file /var/log/app.log.pos + tag rustfs.demo + + @type none + + + + + @type s3 + aws_key_id + aws_sec_key + s3_bucket fluentd-data + s3_endpoint http://rustfs:9000/ + s3_region us-east-1 + force_path_style true + path fluentd-logs + + @type memory + timekey 30s + timekey_wait 0s + flush_mode immediate + + +``` + +`force_path_style true` is required — without it the plugin constructs `fluentd-data.rustfs` as a hostname and every request fails with a DNS error. Inside the Compose network the hostname is `rustfs`; from the host use `http://localhost:9000/`. + +Build the image and start Fluentd on the same Docker network as RustFS: + +```bash +docker build -t fluentd-rustfs . +mkdir -p logs +docker run -d --name fluentd --network oo-rustfs_default \ + -v "$PWD/fluent.conf":/fluentd/etc/fluent.conf:ro \ + -v "$PWD/logs":/var/log fluentd-rustfs +``` + +## 2. Produce log events + +Append lines to the watched file: + +```bash +echo "rustfs fluentd demo line 1" >> logs/app.log +echo "rustfs fluentd demo line 2" >> logs/app.log +``` + +With a 30-second time key and immediate flush mode, each window uploads one gzip object shortly after it closes. Wait about a minute. + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/fluentd-data/ -r +``` + +Each flush window produces one gzipped object: + +```text +fluentd-logs20260921154400_0.gz +fluentd-logs20260921154400_1.gz +``` + +Read one object back to confirm the events are intact: + +```bash +rc cat rustfs/fluentd-data/fluentd-logs20260921154400_0.gz | gunzip +``` + +![Fluentd log objects stored in the RustFS Console](./images/rustfs-fluentd-logs.png) + +## 4. Stop or reset the deployment + +Stop Fluentd while keeping the data: + +```bash +docker rm -f fluentd +``` + +The objects stay in the `fluentd-data` bucket. To delete them, remove the bucket: + +```bash +rc rb rustfs/fluentd-data --force +``` + +## Troubleshooting + +### `Unknown output plugin 's3'` + +The plugin is not installed. Confirm the Dockerfile installs `fluent-plugin-s3` as `root` before switching back to the `fluent` user — installing at container start as the default user fails with permission errors. + +### `Failed to open TCP connection to fluentd-data.rustfs` + +Virtual-hosted addressing is in use. Add `force_path_style true` to the `s3` output so the bucket stays in the URL path. + +### The worker crashes in a restart loop + +The output fails hard when the bucket does not exist. Create `fluentd-data` before starting Fluentd and check the startup log: + +```bash +docker logs fluentd | grep -iE "error|bucket" | tail +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Fluentd outputs. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [fluent-plugin-s3 documentation](https://github.com/fluent/fluent-plugin-s3) for object key formats and compression options. diff --git a/content/en/developer/integration/observability/images/rustfs-fluentd-logs.png b/content/en/developer/integration/observability/images/rustfs-fluentd-logs.png new file mode 100644 index 00000000..3ab768ca Binary files /dev/null and b/content/en/developer/integration/observability/images/rustfs-fluentd-logs.png differ diff --git a/content/en/developer/integration/observability/images/rustfs-otel-logs.png b/content/en/developer/integration/observability/images/rustfs-otel-logs.png new file mode 100644 index 00000000..2f68a22a Binary files /dev/null and b/content/en/developer/integration/observability/images/rustfs-otel-logs.png differ diff --git a/content/en/developer/integration/observability/index.md b/content/en/developer/integration/observability/index.md index dbdbea68..dbce4d0c 100644 --- a/content/en/developer/integration/observability/index.md +++ b/content/en/developer/integration/observability/index.md @@ -7,7 +7,9 @@ Use **RustFS** as the object storage layer for observability platforms that supp ## Platforms +- [Fluentd](./fluentd.md) - [OpenObserve](./openobserve.md) +- [OpenTelemetry](./opentelemetry.md) - [Loki](./loki.md) - [Tempo](./tempo.md) - [Thanos](./thanos.md) diff --git a/content/en/developer/integration/observability/meta.json b/content/en/developer/integration/observability/meta.json index 03f10c67..d6ddebd7 100644 --- a/content/en/developer/integration/observability/meta.json +++ b/content/en/developer/integration/observability/meta.json @@ -1,8 +1,10 @@ { "title": "Observability", "pages": [ - "openobserve", + "fluentd", "loki", + "openobserve", + "opentelemetry", "tempo", "thanos" ] diff --git a/content/en/developer/integration/observability/opentelemetry.md b/content/en/developer/integration/observability/opentelemetry.md new file mode 100644 index 00000000..283df413 --- /dev/null +++ b/content/en/developer/integration/observability/opentelemetry.md @@ -0,0 +1,145 @@ +--- +title: "OpenTelemetry" +description: "Export OpenTelemetry Collector logs to RustFS with the AWS S3 exporter." +--- + +This guide connects the [OpenTelemetry Collector](https://opentelemetry.io/docs/collector/) — the CNCF telemetry pipeline — to **RustFS** through the collector's `awss3` exporter. You will run the contrib collector with a `filelog` receiver, ship the log lines of a file into the `otel-data` bucket, and verify the partitioned objects in RustFS. The workflow was verified with `otel/opentelemetry-collector-contrib:0.138.0` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Log["Application log file"] -->|filelog| Collector["OTel Collector"] + Collector -->|awss3| RustFS["RustFS :9000"] +``` + +The `filelog` receiver tails the log file and the `awss3` exporter uploads batches to the bucket, partitioned into time-based prefixes. Metrics and traces can be routed through the same exporter with their own pipelines. + +## 1. Create the project files + +Create the bucket first — the exporter does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/otel-data +``` + +Create the collector configuration, replacing both credential placeholders: + +```yaml title="config.yaml" +receivers: + filelog: + include: [/var/log/app.log] + start_at: beginning + +exporters: + awss3: + s3uploader: + region: us-east-1 + s3_bucket: otel-data + endpoint: http://rustfs:9000 + s3_force_path_style: true + disable_ssl: true + file_prefix: logs/app + marshaler: body + +service: + pipelines: + logs: + receivers: [filelog] + exporters: [awss3] +``` + +The exporter reads credentials from the standard `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, and `AWS_REGION` environment variables. `marshaler: body` writes each log line as plain text; omit it to store OTLP JSON instead. `s3_force_path_style` and `disable_ssl` are required for a plain-HTTP, non-AWS endpoint. + +Create the Compose file: + +```yaml title="compose.yaml" +services: + collector: + image: otel/opentelemetry-collector-contrib:0.138.0 + command: ["--config=/etc/otelcol-contrib/config.yaml"] + environment: + AWS_ACCESS_KEY_ID: + AWS_SECRET_ACCESS_KEY: + AWS_REGION: us-east-1 + volumes: + - ./config.yaml:/etc/otelcol-contrib/config.yaml:ro + - ./app.log:/var/log/app.log + networks: + - otel + +networks: + otel: +``` + +## 2. Start the collector and produce logs + +Start the stack and append lines to the watched file: + +```bash +echo "otel demo log line one" > app.log +docker compose up -d +echo "second line after start" >> app.log +``` + +The collector tails the file from the beginning and uploads each buffer when the partition rolls over, so allow about a minute after the last line before checking. + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/otel-data/ -r +``` + +Log records are uploaded under time-based partitions with the configured file prefix: + +```text +year=2026/month=09/day=21/hour=15/minute=53/logs/applogs_288361608.txt +``` + +Read one object back to confirm the lines arrived intact: + +```bash +rc cat rustfs/otel-data/year=2026/month=09/day=21/hour=15/minute=53/logs/applogs_288361608.txt +``` + +```text +otel demo log line one +second line after start +``` + +![OpenTelemetry log objects stored in the RustFS Console](./images/rustfs-otel-logs.png) + +## 4. Stop or reset the deployment + +Stop the collector while keeping the data: + +```bash +docker compose down +``` + +The objects stay in the `otel-data` bucket. To delete them, remove the bucket: + +```bash +rc rb rustfs/otel-data --force +``` + +## Troubleshooting + +### `has invalid keys` when the collector starts + +The `awss3` exporter schema differs between collector releases. Version 0.138 nests the upload settings under `s3uploader` as shown above; the endpoint key is `endpoint` (not `s3_endpoint`). Newer releases move these keys to the top level — check the README for your exact collector version. + +### Nothing appears in the bucket + +Confirm the credentials environment variables are set on the collector container, that `s3_force_path_style` is `true`, and that the exporter can reach `http://rustfs:9000` from inside the Compose network. Enable the `debug` exporter on the same pipeline to see whether records flow at all. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional collector operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [AWS S3 exporter documentation](https://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/exporter/awss3exporter) to add metrics and traces pipelines. diff --git a/content/fr/developer/integration/big-data/clickhouse.md b/content/fr/developer/integration/big-data/clickhouse.md new file mode 100644 index 00000000..ac9d8d4b --- /dev/null +++ b/content/fr/developer/integration/big-data/clickhouse.md @@ -0,0 +1,160 @@ +--- +title: "ClickHouse" +description: "Run ClickHouse with an S3 disk backed by RustFS for MergeTree table data." +--- + +This guide connects [ClickHouse](https://github.com/ClickHouse/ClickHouse) — the real-time OLAP database — to **RustFS** through ClickHouse's S3 disk storage policy. You will start a ClickHouse server with Docker, create a MergeTree table that stores its parts on RustFS, insert rows, and verify that the table data lives in the bucket. The workflow was verified with `clickhouse/clickhouse-server:25.8` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["SQL client"] -->|"queries"| CH["ClickHouse :8123"] + CH -->|"MergeTree parts"| RustFS["RustFS :9000"] +``` + +The `rustfs` disk is a ClickHouse S3 disk pointed at the `clickhouse-data` bucket. Tables created with the matching storage policy write their parts — data, index, and checksum files — to the bucket instead of the local filesystem. + +## 1. Create the project files + +Create the bucket first — ClickHouse does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/clickhouse-data +``` + +Create the storage configuration, replacing both credential placeholders: + +```xml title="storage.xml" + + + + + s3 + http://rustfs:9000/clickhouse-data/ + + + + + + + +
+ rustfs +
+
+
+
+
+
+``` + +The endpoint must end with `/` and includes the bucket name as the first path segment. Inside the Compose network the hostname is `rustfs`; from the host use `http://localhost:9000/clickhouse-data/`. + +Start ClickHouse with the configuration mounted: + +```bash +docker run -d --name clickhouse --network oo-rustfs_default \ + -p 8123:8123 \ + -e CLICKHOUSE_PASSWORD= \ + -v "$PWD/storage.xml":/etc/clickhouse-server/config.d/storage.xml:ro \ + clickhouse/clickhouse-server:25.8 +``` + +## 2. Create a table on the S3 disk + +Wait for the HTTP interface, then create a database and a MergeTree table with the storage policy: + +```bash +curl "http://localhost:8123/?password=" \ + --data-binary "CREATE DATABASE rustfs_demo" + +curl "http://localhost:8123/?password=" \ + --data-binary "CREATE TABLE rustfs_demo.events + (id UInt32, name String) + ENGINE = MergeTree ORDER BY id + SETTINGS storage_policy = 'rustfs_policy'" + +curl "http://localhost:8123/?password=" \ + --data-binary "INSERT INTO rustfs_demo.events + VALUES (1, 'clickhouse-on-rustfs'), (2, 'second')" +``` + +Read the rows back and confirm ClickHouse reports the part on the `rustfs` disk: + +```bash +curl "http://localhost:8123/?password=" \ + --data-binary "SELECT count(), any(name) FROM rustfs_demo.events" + +curl "http://localhost:8123/?password=" \ + --data-binary "SELECT name, disk_name FROM system.parts + WHERE database = 'rustfs_demo' AND active" +``` + +```text +2 clickhouse-on-rustfs +all_1_1_0 rustfs +``` + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/clickhouse-data/ -r +``` + +ClickHouse writes each part as content-addressed blobs. The output contains several small objects, and the count grows as more parts are written: + +```text +dtg/hpsyncexixvdnsgorvseobogcgowg +dzp/zfblobhsatzdveqdrsfqcupkdehja +izg/gvhchqobrizpkdftuvlmakkoutwps +``` + +![ClickHouse parts stored in the RustFS Console](./images/rustfs-clickhouse-disk.png) + +The data survives a container restart because the parts live in RustFS: + +```bash +docker restart clickhouse +curl "http://localhost:8123/?password=" \ + --data-binary "SELECT count() FROM rustfs_demo.events" +``` + +## 4. Stop or reset the deployment + +Stop the server while keeping the data: + +```bash +docker rm -f clickhouse +``` + +The parts stay in the `clickhouse-data` bucket and the table can be queried again after the next start. To delete the data, remove the bucket: + +```bash +rc rb rustfs/clickhouse-data --force +``` + +## Troubleshooting + +### `REQUIRED_PASSWORD` on every query + +ClickHouse 25.8 images require a password for the `default` user. Set `CLICKHOUSE_PASSWORD` on the container and pass the same value as the `password` query parameter, as shown above. + +### Table creation fails with a disk or endpoint error + +Confirm that the bucket exists before the table is created, that the endpoint ends with `/`, and that the credentials match the RustFS deployment. Check the server log for the underlying S3 error: + +```bash +docker logs clickhouse | grep -i s3 | tail +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional ClickHouse operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [ClickHouse S3 disk documentation](https://clickhouse.com/docs/engines/table-engines/mergetree-family/mergetree#table_engine-mergetree-s3) to add a cache disk or a tiered hot/cold policy. diff --git a/content/fr/developer/integration/big-data/doris.md b/content/fr/developer/integration/big-data/doris.md new file mode 100644 index 00000000..1c3665b5 --- /dev/null +++ b/content/fr/developer/integration/big-data/doris.md @@ -0,0 +1,161 @@ +--- +title: "Apache Doris" +description: "Back up Apache Doris tables to RustFS through an S3 repository and restore them." +--- + +This guide connects [Apache Doris](https://github.com/apache/doris) — the real-time analytical data warehouse — to **RustFS** through an S3 backup repository. You will start an all-in-one Doris container, create an S3 repository pointing at a RustFS bucket, back up a table, drop it, and restore it from RustFS. The workflow was verified with `apache/doris:all-in-one-4.1.3` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["SQL client"] -->|"queries"| Doris["Doris FE/BE"] + Doris -->|"BACKUP / RESTORE"| RustFS["RustFS :9000"] +``` + +The repository is a named S3 location under the `doris-backups` bucket. `BACKUP SNAPSHOT` uploads table metadata and tablet data files; `RESTORE SNAPSHOT` downloads them into a new table. + +## 1. Start Doris and create the repository + +Create the bucket first — Doris does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/doris-backups +``` + +Start the all-in-one container on the same Docker network as RustFS: + +```bash +docker run -d --name doris --network oo-rustfs_default \ + -p 8030:8030 -p 9030:9030 apache/doris:all-in-one-4.1.3 +``` + +Wait for the frontend to become healthy, then connect with the MySQL protocol (port 9030, user `root`, no password in the all-in-one image). + +Create a test table with rows: + +```sql +CREATE DATABASE rustfs_demo; +CREATE TABLE rustfs_demo.events + (id INT, name VARCHAR(50)) + DISTRIBUTED BY HASH(id) BUCKETS 1 + PROPERTIES ("replication_num" = "1"); +INSERT INTO rustfs_demo.events VALUES (1, 'doris-on-rustfs'), (2, 'backup-test'); +``` + +Create the S3 repository, replacing the endpoint with the IP address of the RustFS container and both credential placeholders: + +```sql +CREATE REPOSITORY `rustfs_repo` + WITH S3 + ON LOCATION "s3://doris-backups/rustfs-repo" + PROPERTIES ( + "AWS_ENDPOINT" = "http://:9000", + "AWS_ACCESS_KEY" = "", + "AWS_SECRET_KEY" = "", + "AWS_REGION" = "us-east-1", + "AWS_PATH_STYLE_ACCESS" = "true" + ); +``` + +Doris 4.1 resolves the bucket into the endpoint hostname even with `AWS_PATH_STYLE_ACCESS` enabled, so a hostname endpoint fails with `UnknownHostException: doris-backups.rustfs`. Using the container IP address forces path-style requests and works; `SHOW REPOSITORIES` confirms the repository registered with an empty `ErrMsg`. + +## 2. Back up a table to RustFS + +Take a snapshot of the table: + +```sql +BACKUP SNAPSHOT rustfs_demo.demo_snapshot + TO rustfs_repo + ON (events); +``` + +The statement returns immediately; the backup job runs in the background. Watch its state: + +```sql +SHOW BACKUP; +``` + +Wait until `State` reaches `FINISHED` — the snapshot metadata and tablet data files are now objects in the bucket. + +## 3. Verify the backup in RustFS + +List the bucket: + +```bash +rc ls rustfs/doris-backups/ -r +``` + +The repository stores a repository descriptor, the snapshot metadata, and the tablet files: + +```text +rustfs-repo/__palo_repository_rustfs_repo/__repo_info +rustfs-repo/__palo_repository_rustfs_repo/__ss_demo_snapshot/__meta.d50ecf9b... +rustfs-repo/__palo_repository_rustfs_repo/__ss_demo_snapshot/__ss_content/.../...dat... +``` + +![Doris backup objects stored in the RustFS Console](./images/rustfs-doris-backup.png) + +## 4. Restore the table from RustFS + +Drop the table and restore it from the snapshot. The timestamp comes from the snapshot name shown by `SHOW SNAPSHOT ON REPOSITORY rustfs_repo;`: + +```sql +DROP TABLE rustfs_demo.events; + +RESTORE SNAPSHOT rustfs_demo.demo_snapshot + FROM rustfs_repo + ON (events) + PROPERTIES ( + "backup_timestamp" = "2026-09-21-16-35-12", + "replication_num" = "1" + ); +``` + +Wait for the restore job to finish and confirm the data: + +```sql +SHOW RESTORE; +SELECT count(*) FROM rustfs_demo.events; +``` + +```text +2 +``` + +## 5. Stop or reset the deployment + +Stop Doris while keeping the data: + +```bash +docker rm -f doris +``` + +The backup stays in the `doris-backups` bucket and can be restored into any Doris cluster that registers the same repository. To delete it, remove the bucket: + +```bash +rc rb rustfs/doris-backups --force +``` + +## Troubleshooting + +### `UnknownHostException: doris-backups.rustfs` when creating the repository + +Doris is building a virtual-hosted hostname from the bucket and endpoint. Use the RustFS container IP address in `AWS_ENDPOINT` together with `AWS_PATH_STYLE_ACCESS = "true"`, as shown above. + +### The backup stays in `SNAPSHOTING` for a long time + +The backend uploads the tablet files. Confirm the backend is healthy (`SHOW BACKENDS;`) and can reach the endpoint; the all-in-one image needs a minute or two after start before both processes report ready. + +### `Failed to create repository: ... file status` + +The bucket does not exist or the credentials are wrong. Create `doris-backups` with `rc mb` and re-check the access key pair. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Doris operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Doris backup and restore documentation](https://doris.apache.org/docs/data-operate/backup-restore/) to schedule periodic snapshots. diff --git a/content/fr/developer/integration/big-data/images/rustfs-clickhouse-disk.png b/content/fr/developer/integration/big-data/images/rustfs-clickhouse-disk.png new file mode 100644 index 00000000..9225bf66 Binary files /dev/null and b/content/fr/developer/integration/big-data/images/rustfs-clickhouse-disk.png differ diff --git a/content/fr/developer/integration/big-data/images/rustfs-doris-backup.png b/content/fr/developer/integration/big-data/images/rustfs-doris-backup.png new file mode 100644 index 00000000..9a936276 Binary files /dev/null and b/content/fr/developer/integration/big-data/images/rustfs-doris-backup.png differ diff --git a/content/fr/developer/integration/big-data/images/rustfs-opendal-objects.png b/content/fr/developer/integration/big-data/images/rustfs-opendal-objects.png new file mode 100644 index 00000000..d7129423 Binary files /dev/null and b/content/fr/developer/integration/big-data/images/rustfs-opendal-objects.png differ diff --git a/content/fr/developer/integration/big-data/images/rustfs-zeppelin-notebook.png b/content/fr/developer/integration/big-data/images/rustfs-zeppelin-notebook.png new file mode 100644 index 00000000..9c5fe782 Binary files /dev/null and b/content/fr/developer/integration/big-data/images/rustfs-zeppelin-notebook.png differ diff --git a/content/fr/developer/integration/big-data/index.md b/content/fr/developer/integration/big-data/index.md index e0d11e44..41f295dd 100644 --- a/content/fr/developer/integration/big-data/index.md +++ b/content/fr/developer/integration/big-data/index.md @@ -7,11 +7,14 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo ## Systems +- [ClickHouse](./clickhouse.md) - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) - [Milvus](./milvus.md) - [MLflow](./mlflow.md) +- [OpenDAL](./opendal.md) - [DuckDB](./duckdb.md) +- [Doris](./doris.md) - [InfluxDB](./influxdb.md) - [Spark](./spark.md) - [Flink](./flink.md) diff --git a/content/fr/developer/integration/big-data/meta.json b/content/fr/developer/integration/big-data/meta.json index 87e29f91..ee7d69b5 100644 --- a/content/fr/developer/integration/big-data/meta.json +++ b/content/fr/developer/integration/big-data/meta.json @@ -1,14 +1,18 @@ { "title": "Data Analytics", "pages": [ + "clickhouse", "iceberg", "pyiceberg", "milvus", "mlflow", + "opendal", "duckdb", + "doris", "influxdb", "spark", "flink", - "trino" + "trino", + "zeppelin" ] } diff --git a/content/fr/developer/integration/big-data/opendal.md b/content/fr/developer/integration/big-data/opendal.md new file mode 100644 index 00000000..074b7475 --- /dev/null +++ b/content/fr/developer/integration/big-data/opendal.md @@ -0,0 +1,108 @@ +--- +title: "OpenDAL" +description: "Access RustFS objects from applications through the Apache OpenDAL data access layer." +--- + +This guide connects [Apache OpenDAL](https://github.com/apache/opendal) — the unified data access layer — to **RustFS** through its `s3` service. You will run the OpenDAL Python binding against RustFS, write and read an object, list a prefix, and delete it. The workflow was verified with the `opendal` Python package 0.46 on `python:3.12-slim` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and Python 3.9 or later. OpenDAL supports the same `s3` service from Rust, Java, Node.js, and Go bindings with equivalent settings. + +## Architecture + +```mermaid +flowchart LR + App["Application"] -->|"Operator API"| OpenDAL["OpenDAL"] + OpenDAL -->|"s3 service"| RustFS["RustFS :9000"] +``` + +An OpenDAL `Operator` bound to the `s3` service exposes one uniform API — `write`, `read`, `stat`, `list`, `delete` — over the bucket, so the same code runs against S3, RustFS, or any other supported service by changing the connection settings. + +## 1. Set up the project + +Install the Python binding: + +```bash +pip install opendal +``` + +Create the script, replacing all connection placeholders: + +```python title="opendal_demo.py" +import opendal + +op = opendal.Operator( + "s3", + endpoint="http://:9000", + bucket="my-bucket", + access_key_id="", + secret_access_key="", + region="us-east-1", +) +op.write("opendal-demo/hello.txt", b"hello from opendal against rustfs") +print("read-back:", op.read("opendal-demo/hello.txt")) +print("content_length:", op.stat("opendal-demo/hello.txt").content_length) +for entry in op.list("opendal-demo/"): + print("listed:", entry.path) +op.delete("opendal-demo/hello.txt") +print("deleted:", not op.exists("opendal-demo/hello.txt")) +``` + +The credential option names are `access_key_id` and `secret_access_key` — the shorter `access_key` names do not exist and fail with a signing error. Path-style addressing is the default for non-AWS endpoints. + +## 2. Run the demo + +Run the script from a machine that can reach RustFS: + +```bash +python opendal_demo.py +``` + +```text +read-back: b"hello from opendal against rustfs" +content_length: 33 +listed: opendal-demo/hello.txt +deleted: True +``` + +The round trip exercises the full object lifecycle: write uploads the bytes, `read` fetches them back, `stat` returns the object size, `list` enumerates the prefix, and `delete` removes the object. + +## 3. Verify objects in RustFS + +Comment out the final `op.delete` line, run the script again, and list the prefix in RustFS: + +```bash +rc ls rustfs/my-bucket/opendal-demo/ -r +``` + +```text +hello.txt +data/rows.csv +``` + +![OpenDAL objects stored in the RustFS Console](./images/rustfs-opendal-objects.png) + +The objects visible in the RustFS Console are exactly the paths the OpenDAL API wrote. + +## 4. Stop or reset + +OpenDAL is a library and holds no state of its own. To clean up the demo objects: + +```bash +rc rm rustfs/my-bucket/opendal-demo/ --recursive --force +``` + +## Troubleshooting + +### `failed to load signing credential` + +The operator received no usable credentials. Use the exact option names `access_key_id` and `secret_access_key`; other spellings are silently ignored and signing then fails. + +### Connection or DNS errors on write + +Confirm the endpoint includes the scheme and port and is reachable from the application. Inside a Compose network the hostname is `rustfs`; from the host use `http://localhost:9000`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional OpenDAL operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [OpenDAL documentation](https://opendal.apache.org/docs/) to use the same operator from Rust, Java, or Node.js. diff --git a/content/fr/developer/integration/big-data/zeppelin.md b/content/fr/developer/integration/big-data/zeppelin.md new file mode 100644 index 00000000..e7f47375 --- /dev/null +++ b/content/fr/developer/integration/big-data/zeppelin.md @@ -0,0 +1,109 @@ +--- +title: "Apache Zeppelin" +description: "Store Apache Zeppelin notebooks in RustFS through the S3 notebook repository." +--- + +This guide connects [Apache Zeppelin](https://github.com/apache/zeppelin) — the web-based notebook for data analytics — to **RustFS** through Zeppelin's S3 notebook storage. You will start Zeppelin pointed at a RustFS bucket, create a note, and verify that the notebook file is stored in RustFS and survives a restart. The workflow was verified with `apache/zeppelin:0.12.0` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Browser["Browser"] -->|"notebook edits"| Z["Zeppelin :8080"] + Z -->|".zpln files"| RustFS["RustFS :9000"] +``` + +With the `S3NotebookRepo` storage class, every note is persisted as a `.zpln` JSON file under the `user/notebook/` prefix of the bucket. Zeppelin reads and writes the bucket directly, so notes survive container restarts and can be shared between instances. + +## 1. Start Zeppelin + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +AWS_ACCESS_KEY_ID= +AWS_SECRET_ACCESS_KEY= +``` + +Start Zeppelin on the same Docker network as RustFS with the S3 storage settings: + +```bash +docker run -d --name zeppelin --network oo-rustfs_default \ + -p 8080:8080 \ + -e AWS_ACCESS_KEY_ID \ + -e AWS_SECRET_ACCESS_KEY \ + -e ZEPPELIN_NOTEBOOK_STORAGE=org.apache.zeppelin.notebook.repo.S3NotebookRepo \ + -e ZEPPELIN_NOTEBOOK_S3_BUCKET=my-bucket \ + -e ZEPPELIN_NOTEBOOK_S3_ENDPOINT=http://rustfs:9000 \ + -e ZEPPELIN_NOTEBOOK_S3_PATH_STYLE_ACCESS=true \ + apache/zeppelin:0.12.0 +``` + +Zeppelin reads `ZEPPELIN_*` environment variables as configuration properties, so no `zeppelin-site.xml` edit is needed. This guide uses the existing `my-bucket`; notes land under its `user/notebook/` prefix, which the S3 storage creates on demand. + +## 2. Create a note + +Wait for the UI on `http://localhost:8080`, then create a note named `rustfs-demo` in the notebook list and add a paragraph, or use the REST API: + +```bash +NOTE=$(curl -s -X POST "http://localhost:8080/api/notebook" \ + -H "Content-Type: application/json" \ + -d '{"name": "rustfs-demo"}' | python3 -c "import json,sys; print(json.load(sys.stdin)['body'])") +echo "note id: $NOTE" +``` + +## 3. Verify the notebook in RustFS + +List the notebook prefix: + +```bash +rc ls rustfs/my-bucket/user/notebook/ -r +``` + +The note is stored as a JSON file named after the note and its ID: + +```text +user/notebook/rustfs-demo_2N4PY7UY5.zpln +``` + +![Zeppelin notebooks stored in the RustFS Console](./images/rustfs-zeppelin-notebook.png) + +Notes survive a restart because Zeppelin loads them from the bucket: + +```bash +docker restart zeppelin +curl -s "http://localhost:8080/api/notebook" | head -c 200 +``` + +The note list contains `2N4PY7UY5` again after the restart. + +## 4. Stop or reset the deployment + +Stop Zeppelin while keeping the notes: + +```bash +docker rm -f zeppelin +``` + +The notes stay in the `user/notebook/` prefix of `my-bucket`. To delete them, remove the prefix: + +```bash +rc rm rustfs/my-bucket/user/notebook/ --recursive --force +``` + +## Troubleshooting + +### Zeppelin starts but notes never appear in the bucket + +Confirm the three `ZEPPELIN_NOTEBOOK_S3_*` variables are set and that the credentials environment variables reach the container — the S3 repository is initialized at startup, so the container must be recreated after any change. + +### `UnknownHostException: my-bucket.rustfs` + +The path-style flag was not picked up. Keep `ZEPPELIN_NOTEBOOK_S3_PATH_STYLE_ACCESS=true` exactly as shown; with path-style disabled Zeppelin treats the bucket as a hostname. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Zeppelin operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Zeppelin storage documentation](https://zeppelin.apache.org/docs/latest/setup/storage/storage.html#notebook-storage-in-s3) to organize notes in per-user prefixes. diff --git a/content/fr/developer/integration/devops/images/rustfs-jenkins-artifacts.png b/content/fr/developer/integration/devops/images/rustfs-jenkins-artifacts.png new file mode 100644 index 00000000..19c8cf87 Binary files /dev/null and b/content/fr/developer/integration/devops/images/rustfs-jenkins-artifacts.png differ diff --git a/content/fr/developer/integration/devops/index.md b/content/fr/developer/integration/devops/index.md index 0f991539..b66e4aa0 100644 --- a/content/fr/developer/integration/devops/index.md +++ b/content/fr/developer/integration/devops/index.md @@ -9,6 +9,7 @@ Utilisez **RustFS** comme couche de stockage objet pour les plateformes DevOps e - [Elasticsearch](./elasticsearch.md) - [Gitea](./gitea.md) +- [Jenkins](./jenkins.md) - [Terraform](./terraform.md) Conservez les artefacts, l'état et les données de télémétrie dans des buckets dédiés et limitez les identifiants aux opérations de bucket requises. diff --git a/content/fr/developer/integration/devops/jenkins.md b/content/fr/developer/integration/devops/jenkins.md new file mode 100644 index 00000000..1d78f31b --- /dev/null +++ b/content/fr/developer/integration/devops/jenkins.md @@ -0,0 +1,144 @@ +--- +title: "Jenkins" +description: "Store Jenkins build artifacts in RustFS with the Artifact Manager on S3 plugin." +--- + +This guide connects [Jenkins](https://github.com/jenkinsci/jenkins) — the automation server — to **RustFS** through the Artifact Manager on S3 plugin. You will start Jenkins with the plugin, point its artifact manager at a RustFS bucket, run a job that archives an artifact, and verify that the artifact is stored in RustFS. The workflow was verified with `jenkins/jenkins:lts-jdk17` (Jenkins 2.5xx LTS) and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Dev["Developer"] -->|"trigger build"| J["Jenkins :8080"] + J -->|"archiveArtifacts"| RustFS["RustFS :9000"] +``` + +With the plugin active, every job that publishes artifacts through the standard `archiveArtifacts` step (or `stash`/`unstash`) uploads them to the `jenkins-artifacts` bucket under the configured prefix instead of storing them on the controller disk. + +## 1. Create the project files + +Create the bucket first — the plugin validates the bucket but does not create it: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/jenkins-artifacts +``` + +Create a Dockerfile that installs the plugin into the Jenkins LTS image: + +```dockerfile title="Dockerfile" +FROM jenkins/jenkins:lts-jdk17 +USER root +RUN jenkins-plugin-cli --plugins artifact-manager-s3 aws-credentials +USER jenkins +``` + +The `artifact-manager-s3` plugin brings in the AWS credentials support it needs; listing `aws-credentials` explicitly keeps the credential type available. + +Build and start Jenkins on the same Docker network as RustFS: + +```bash +docker build -t jenkins-rustfs . +docker run -d --name jenkins --network oo-rustfs_default \ + -p 8080:8080 -v jenkins-home:/var/jenkins_home jenkins-rustfs +``` + +Complete the setup wizard, then create an AWS credential: **Manage Jenkins → Credentials → global → Add Credentials**, kind **AWS Credential**, ID `rustfs-creds`, with your RustFS access key and secret key. + +## 2. Configure the artifact manager + +Open **Manage Jenkins → AWS Configuration** (from the `aws-global-configuration` plugin) and set: + +- **Region name**: `us-east-1` +- **Credentials**: `rustfs-creds` + +Open **Manage Jenkins → System** and locate the **Artifact Management for Builds** section. Select **Cloud Provider Amazon S3**, and fill in the S3 configuration: + +- **S3 Bucket Name**: `jenkins-artifacts` +- **S3 Bucket Region**: `us-east-1` +- **Base Prefix**: `artifacts/` +- **Custom Endpoint**: `:9000` (for example `rustfs:9000` inside the Compose network) +- **Custom Signing Region**: `us-east-1` +- **Use Path Style URL**: enabled +- **Use Insecure HTTP**: enabled +- **Disable Session Token**: enabled + +Path-style addressing and plain HTTP are required for a non-AWS endpoint without TLS. Disabling the session token stops the plugin from calling AWS STS, which a plain access key pair cannot answer. Click **Validate S3 Bucket configuration** to confirm the settings, then save. + +## 3. Run a job that archives an artifact + +Create a freestyle job (or pipeline) that produces a file and archives it: + +```groovy +pipeline { + agent any + stages { + stage('Build') { + steps { + sh 'echo "jenkins artifact stored on rustfs" > report.txt' + } + } + } + post { + always { + archiveArtifacts 'report.txt' + } + } +} +``` + +Run the build and wait for it to finish. The artifact upload goes to RustFS transparently — the job configuration does not mention S3 at all. + +## 4. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/jenkins-artifacts/ -r +``` + +The artifact is stored under the prefix, organized by job and build number: + +```text +artifacts/s3-artifacts-demo/3/artifacts/report.txt +``` + +![Jenkins artifacts stored in the RustFS Console](./images/rustfs-jenkins-artifacts.png) + +Downloading the artifact from the build page reads it back from RustFS. + +## 5. Stop or reset the deployment + +Stop Jenkins while keeping the data: + +```bash +docker rm -f jenkins +``` + +The artifacts stay in the `jenkins-artifacts` bucket. To delete them, remove the bucket: + +```bash +rc rb rustfs/jenkins-artifacts --force +``` + +## Troubleshooting + +### `StsException: The security token included in the request is invalid` + +The plugin is calling AWS STS to obtain session credentials. Enable **Disable Session Token** in the S3 configuration — a static access key pair cannot answer an STS call. + +### `UnknownHostException: jenkins-artifacts.rustfs` + +The plugin is using virtual-hosted addressing. Enable **Use Path Style URL** — RustFS resolves buckets from the URL path, not the hostname. + +### `No valid session credentials` or empty credential errors + +Confirm the AWS Configuration page has the credential selected and saved **before** the S3 bucket settings are used, and that the credential ID matches the one you created. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Jenkins integrations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Artifact Manager on S3 plugin documentation](https://plugins.jenkins.io/artifact-manager-s3/) for stash support and cleanup options. diff --git a/content/fr/developer/integration/devops/meta.json b/content/fr/developer/integration/devops/meta.json index 61ad9fb8..67c84c6e 100644 --- a/content/fr/developer/integration/devops/meta.json +++ b/content/fr/developer/integration/devops/meta.json @@ -3,6 +3,7 @@ "pages": [ "elasticsearch", "gitea", + "jenkins", "terraform" ] } diff --git a/content/fr/developer/integration/index.md b/content/fr/developer/integration/index.md index 0af98703..c7404168 100644 --- a/content/fr/developer/integration/index.md +++ b/content/fr/developer/integration/index.md @@ -9,10 +9,10 @@ Utilisez cette section pour connecter **RustFS** à des plateformes d'infrastruc - [Reverse Proxy](./reverse-proxy/index.md) couvre Nginx, Traefik, Caddy et HAProxy. - [Backup](./backup/index.md) couvre Restic et Longhorn. -- [Analyse de données](./big-data/index.md) couvre Iceberg. -- [Observabilité](./observability/index.md) couvre OpenObserve. +- [Analyse de données](./big-data/index.md) couvre les systèmes d'analyse incluant ClickHouse, Doris, Iceberg, Milvus, OpenDAL et Zeppelin. +- [Observabilité](./observability/index.md) couvre les systèmes de télémétrie incluant Fluentd, OpenObserve, OpenTelemetry, Thanos et Tempo. - [Autres](./others/index.md) couvre le SDK communautaire capo pour Python. - [Registre](./registry/index.md) couvre Harbor. -- [DevOps](./devops/index.md) couvre Elasticsearch, Gitea et Terraform. +- [DevOps](./devops/index.md) couvre Elasticsearch, Gitea, Jenkins et Terraform. Chaque guide indique le point de terminaison RustFS et les exigences d'adressage à utiliser lors de la configuration du système intégré. \ No newline at end of file diff --git a/content/fr/developer/integration/observability/fluentd.md b/content/fr/developer/integration/observability/fluentd.md new file mode 100644 index 00000000..4607c0ec --- /dev/null +++ b/content/fr/developer/integration/observability/fluentd.md @@ -0,0 +1,152 @@ +--- +title: "Fluentd" +description: "Ship Fluentd log events to RustFS with the S3 output plugin." +--- + +This guide connects [Fluentd](https://github.com/fluent/fluentd) — the open-source data collector — to **RustFS** through the `out_s3` output plugin. You will run Fluentd with a tail source, buffer log events, and verify that the flushed objects are stored in RustFS. The workflow was verified with `fluent/fluentd:v1.17-1`, `fluent-plugin-s3` 1.8.6, and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + App["Application"] -->|"writes lines"| File["app.log"] + File -->|tail| Fluentd["Fluentd"] + Fluentd -->|"gzip objects"| RustFS["RustFS :9000"] +``` + +The tail source reads new lines from `app.log` and hands them to the S3 output, which buffers events on a time key and uploads a gzip object per flush window. + +## 1. Create the project files + +Create the bucket first — Fluentd does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/fluentd-data +``` + +The official Fluentd image runs as a non-root user and cannot install gems at startup, so build a small image with the S3 plugin: + +```dockerfile title="Dockerfile" +FROM fluent/fluentd:v1.17-1 +USER root +RUN gem install fluent-plugin-s3 --no-document +USER fluent +``` + +Create the Fluentd configuration, replacing both credential placeholders: + +```nginx title="fluent.conf" + + @type tail + path /var/log/app.log + pos_file /var/log/app.log.pos + tag rustfs.demo + + @type none + + + + + @type s3 + aws_key_id + aws_sec_key + s3_bucket fluentd-data + s3_endpoint http://rustfs:9000/ + s3_region us-east-1 + force_path_style true + path fluentd-logs + + @type memory + timekey 30s + timekey_wait 0s + flush_mode immediate + + +``` + +`force_path_style true` is required — without it the plugin constructs `fluentd-data.rustfs` as a hostname and every request fails with a DNS error. Inside the Compose network the hostname is `rustfs`; from the host use `http://localhost:9000/`. + +Build the image and start Fluentd on the same Docker network as RustFS: + +```bash +docker build -t fluentd-rustfs . +mkdir -p logs +docker run -d --name fluentd --network oo-rustfs_default \ + -v "$PWD/fluent.conf":/fluentd/etc/fluent.conf:ro \ + -v "$PWD/logs":/var/log fluentd-rustfs +``` + +## 2. Produce log events + +Append lines to the watched file: + +```bash +echo "rustfs fluentd demo line 1" >> logs/app.log +echo "rustfs fluentd demo line 2" >> logs/app.log +``` + +With a 30-second time key and immediate flush mode, each window uploads one gzip object shortly after it closes. Wait about a minute. + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/fluentd-data/ -r +``` + +Each flush window produces one gzipped object: + +```text +fluentd-logs20260921154400_0.gz +fluentd-logs20260921154400_1.gz +``` + +Read one object back to confirm the events are intact: + +```bash +rc cat rustfs/fluentd-data/fluentd-logs20260921154400_0.gz | gunzip +``` + +![Fluentd log objects stored in the RustFS Console](./images/rustfs-fluentd-logs.png) + +## 4. Stop or reset the deployment + +Stop Fluentd while keeping the data: + +```bash +docker rm -f fluentd +``` + +The objects stay in the `fluentd-data` bucket. To delete them, remove the bucket: + +```bash +rc rb rustfs/fluentd-data --force +``` + +## Troubleshooting + +### `Unknown output plugin 's3'` + +The plugin is not installed. Confirm the Dockerfile installs `fluent-plugin-s3` as `root` before switching back to the `fluent` user — installing at container start as the default user fails with permission errors. + +### `Failed to open TCP connection to fluentd-data.rustfs` + +Virtual-hosted addressing is in use. Add `force_path_style true` to the `s3` output so the bucket stays in the URL path. + +### The worker crashes in a restart loop + +The output fails hard when the bucket does not exist. Create `fluentd-data` before starting Fluentd and check the startup log: + +```bash +docker logs fluentd | grep -iE "error|bucket" | tail +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Fluentd outputs. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [fluent-plugin-s3 documentation](https://github.com/fluent/fluent-plugin-s3) for object key formats and compression options. diff --git a/content/fr/developer/integration/observability/images/rustfs-fluentd-logs.png b/content/fr/developer/integration/observability/images/rustfs-fluentd-logs.png new file mode 100644 index 00000000..3ab768ca Binary files /dev/null and b/content/fr/developer/integration/observability/images/rustfs-fluentd-logs.png differ diff --git a/content/fr/developer/integration/observability/images/rustfs-otel-logs.png b/content/fr/developer/integration/observability/images/rustfs-otel-logs.png new file mode 100644 index 00000000..2f68a22a Binary files /dev/null and b/content/fr/developer/integration/observability/images/rustfs-otel-logs.png differ diff --git a/content/fr/developer/integration/observability/index.md b/content/fr/developer/integration/observability/index.md index 2094641a..8850eba3 100644 --- a/content/fr/developer/integration/observability/index.md +++ b/content/fr/developer/integration/observability/index.md @@ -7,7 +7,9 @@ Utilisez **RustFS** comme couche de stockage objet pour les plateformes d'observ ## Plateformes +- [Fluentd](./fluentd.md) - [OpenObserve](./openobserve.md) +- [OpenTelemetry](./opentelemetry.md) - [Loki](./loki.md) - [Tempo](./tempo.md) - [Thanos](./thanos.md) diff --git a/content/fr/developer/integration/observability/meta.json b/content/fr/developer/integration/observability/meta.json index 1709c8c7..85efe3c7 100644 --- a/content/fr/developer/integration/observability/meta.json +++ b/content/fr/developer/integration/observability/meta.json @@ -1,8 +1,10 @@ { "title": "Observabilité", "pages": [ - "openobserve", + "fluentd", "loki", + "openobserve", + "opentelemetry", "tempo", "thanos" ] diff --git a/content/fr/developer/integration/observability/opentelemetry.md b/content/fr/developer/integration/observability/opentelemetry.md new file mode 100644 index 00000000..283df413 --- /dev/null +++ b/content/fr/developer/integration/observability/opentelemetry.md @@ -0,0 +1,145 @@ +--- +title: "OpenTelemetry" +description: "Export OpenTelemetry Collector logs to RustFS with the AWS S3 exporter." +--- + +This guide connects the [OpenTelemetry Collector](https://opentelemetry.io/docs/collector/) — the CNCF telemetry pipeline — to **RustFS** through the collector's `awss3` exporter. You will run the contrib collector with a `filelog` receiver, ship the log lines of a file into the `otel-data` bucket, and verify the partitioned objects in RustFS. The workflow was verified with `otel/opentelemetry-collector-contrib:0.138.0` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Log["Application log file"] -->|filelog| Collector["OTel Collector"] + Collector -->|awss3| RustFS["RustFS :9000"] +``` + +The `filelog` receiver tails the log file and the `awss3` exporter uploads batches to the bucket, partitioned into time-based prefixes. Metrics and traces can be routed through the same exporter with their own pipelines. + +## 1. Create the project files + +Create the bucket first — the exporter does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/otel-data +``` + +Create the collector configuration, replacing both credential placeholders: + +```yaml title="config.yaml" +receivers: + filelog: + include: [/var/log/app.log] + start_at: beginning + +exporters: + awss3: + s3uploader: + region: us-east-1 + s3_bucket: otel-data + endpoint: http://rustfs:9000 + s3_force_path_style: true + disable_ssl: true + file_prefix: logs/app + marshaler: body + +service: + pipelines: + logs: + receivers: [filelog] + exporters: [awss3] +``` + +The exporter reads credentials from the standard `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, and `AWS_REGION` environment variables. `marshaler: body` writes each log line as plain text; omit it to store OTLP JSON instead. `s3_force_path_style` and `disable_ssl` are required for a plain-HTTP, non-AWS endpoint. + +Create the Compose file: + +```yaml title="compose.yaml" +services: + collector: + image: otel/opentelemetry-collector-contrib:0.138.0 + command: ["--config=/etc/otelcol-contrib/config.yaml"] + environment: + AWS_ACCESS_KEY_ID: + AWS_SECRET_ACCESS_KEY: + AWS_REGION: us-east-1 + volumes: + - ./config.yaml:/etc/otelcol-contrib/config.yaml:ro + - ./app.log:/var/log/app.log + networks: + - otel + +networks: + otel: +``` + +## 2. Start the collector and produce logs + +Start the stack and append lines to the watched file: + +```bash +echo "otel demo log line one" > app.log +docker compose up -d +echo "second line after start" >> app.log +``` + +The collector tails the file from the beginning and uploads each buffer when the partition rolls over, so allow about a minute after the last line before checking. + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/otel-data/ -r +``` + +Log records are uploaded under time-based partitions with the configured file prefix: + +```text +year=2026/month=09/day=21/hour=15/minute=53/logs/applogs_288361608.txt +``` + +Read one object back to confirm the lines arrived intact: + +```bash +rc cat rustfs/otel-data/year=2026/month=09/day=21/hour=15/minute=53/logs/applogs_288361608.txt +``` + +```text +otel demo log line one +second line after start +``` + +![OpenTelemetry log objects stored in the RustFS Console](./images/rustfs-otel-logs.png) + +## 4. Stop or reset the deployment + +Stop the collector while keeping the data: + +```bash +docker compose down +``` + +The objects stay in the `otel-data` bucket. To delete them, remove the bucket: + +```bash +rc rb rustfs/otel-data --force +``` + +## Troubleshooting + +### `has invalid keys` when the collector starts + +The `awss3` exporter schema differs between collector releases. Version 0.138 nests the upload settings under `s3uploader` as shown above; the endpoint key is `endpoint` (not `s3_endpoint`). Newer releases move these keys to the top level — check the README for your exact collector version. + +### Nothing appears in the bucket + +Confirm the credentials environment variables are set on the collector container, that `s3_force_path_style` is `true`, and that the exporter can reach `http://rustfs:9000` from inside the Compose network. Enable the `debug` exporter on the same pipeline to see whether records flow at all. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional collector operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [AWS S3 exporter documentation](https://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/exporter/awss3exporter) to add metrics and traces pipelines. diff --git a/content/ja/developer/integration/big-data/clickhouse.md b/content/ja/developer/integration/big-data/clickhouse.md new file mode 100644 index 00000000..ac9d8d4b --- /dev/null +++ b/content/ja/developer/integration/big-data/clickhouse.md @@ -0,0 +1,160 @@ +--- +title: "ClickHouse" +description: "Run ClickHouse with an S3 disk backed by RustFS for MergeTree table data." +--- + +This guide connects [ClickHouse](https://github.com/ClickHouse/ClickHouse) — the real-time OLAP database — to **RustFS** through ClickHouse's S3 disk storage policy. You will start a ClickHouse server with Docker, create a MergeTree table that stores its parts on RustFS, insert rows, and verify that the table data lives in the bucket. The workflow was verified with `clickhouse/clickhouse-server:25.8` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["SQL client"] -->|"queries"| CH["ClickHouse :8123"] + CH -->|"MergeTree parts"| RustFS["RustFS :9000"] +``` + +The `rustfs` disk is a ClickHouse S3 disk pointed at the `clickhouse-data` bucket. Tables created with the matching storage policy write their parts — data, index, and checksum files — to the bucket instead of the local filesystem. + +## 1. Create the project files + +Create the bucket first — ClickHouse does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/clickhouse-data +``` + +Create the storage configuration, replacing both credential placeholders: + +```xml title="storage.xml" + + + + + s3 + http://rustfs:9000/clickhouse-data/ + + + + + + + +
+ rustfs +
+
+
+
+
+
+``` + +The endpoint must end with `/` and includes the bucket name as the first path segment. Inside the Compose network the hostname is `rustfs`; from the host use `http://localhost:9000/clickhouse-data/`. + +Start ClickHouse with the configuration mounted: + +```bash +docker run -d --name clickhouse --network oo-rustfs_default \ + -p 8123:8123 \ + -e CLICKHOUSE_PASSWORD= \ + -v "$PWD/storage.xml":/etc/clickhouse-server/config.d/storage.xml:ro \ + clickhouse/clickhouse-server:25.8 +``` + +## 2. Create a table on the S3 disk + +Wait for the HTTP interface, then create a database and a MergeTree table with the storage policy: + +```bash +curl "http://localhost:8123/?password=" \ + --data-binary "CREATE DATABASE rustfs_demo" + +curl "http://localhost:8123/?password=" \ + --data-binary "CREATE TABLE rustfs_demo.events + (id UInt32, name String) + ENGINE = MergeTree ORDER BY id + SETTINGS storage_policy = 'rustfs_policy'" + +curl "http://localhost:8123/?password=" \ + --data-binary "INSERT INTO rustfs_demo.events + VALUES (1, 'clickhouse-on-rustfs'), (2, 'second')" +``` + +Read the rows back and confirm ClickHouse reports the part on the `rustfs` disk: + +```bash +curl "http://localhost:8123/?password=" \ + --data-binary "SELECT count(), any(name) FROM rustfs_demo.events" + +curl "http://localhost:8123/?password=" \ + --data-binary "SELECT name, disk_name FROM system.parts + WHERE database = 'rustfs_demo' AND active" +``` + +```text +2 clickhouse-on-rustfs +all_1_1_0 rustfs +``` + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/clickhouse-data/ -r +``` + +ClickHouse writes each part as content-addressed blobs. The output contains several small objects, and the count grows as more parts are written: + +```text +dtg/hpsyncexixvdnsgorvseobogcgowg +dzp/zfblobhsatzdveqdrsfqcupkdehja +izg/gvhchqobrizpkdftuvlmakkoutwps +``` + +![ClickHouse parts stored in the RustFS Console](./images/rustfs-clickhouse-disk.png) + +The data survives a container restart because the parts live in RustFS: + +```bash +docker restart clickhouse +curl "http://localhost:8123/?password=" \ + --data-binary "SELECT count() FROM rustfs_demo.events" +``` + +## 4. Stop or reset the deployment + +Stop the server while keeping the data: + +```bash +docker rm -f clickhouse +``` + +The parts stay in the `clickhouse-data` bucket and the table can be queried again after the next start. To delete the data, remove the bucket: + +```bash +rc rb rustfs/clickhouse-data --force +``` + +## Troubleshooting + +### `REQUIRED_PASSWORD` on every query + +ClickHouse 25.8 images require a password for the `default` user. Set `CLICKHOUSE_PASSWORD` on the container and pass the same value as the `password` query parameter, as shown above. + +### Table creation fails with a disk or endpoint error + +Confirm that the bucket exists before the table is created, that the endpoint ends with `/`, and that the credentials match the RustFS deployment. Check the server log for the underlying S3 error: + +```bash +docker logs clickhouse | grep -i s3 | tail +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional ClickHouse operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [ClickHouse S3 disk documentation](https://clickhouse.com/docs/engines/table-engines/mergetree-family/mergetree#table_engine-mergetree-s3) to add a cache disk or a tiered hot/cold policy. diff --git a/content/ja/developer/integration/big-data/doris.md b/content/ja/developer/integration/big-data/doris.md new file mode 100644 index 00000000..1c3665b5 --- /dev/null +++ b/content/ja/developer/integration/big-data/doris.md @@ -0,0 +1,161 @@ +--- +title: "Apache Doris" +description: "Back up Apache Doris tables to RustFS through an S3 repository and restore them." +--- + +This guide connects [Apache Doris](https://github.com/apache/doris) — the real-time analytical data warehouse — to **RustFS** through an S3 backup repository. You will start an all-in-one Doris container, create an S3 repository pointing at a RustFS bucket, back up a table, drop it, and restore it from RustFS. The workflow was verified with `apache/doris:all-in-one-4.1.3` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["SQL client"] -->|"queries"| Doris["Doris FE/BE"] + Doris -->|"BACKUP / RESTORE"| RustFS["RustFS :9000"] +``` + +The repository is a named S3 location under the `doris-backups` bucket. `BACKUP SNAPSHOT` uploads table metadata and tablet data files; `RESTORE SNAPSHOT` downloads them into a new table. + +## 1. Start Doris and create the repository + +Create the bucket first — Doris does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/doris-backups +``` + +Start the all-in-one container on the same Docker network as RustFS: + +```bash +docker run -d --name doris --network oo-rustfs_default \ + -p 8030:8030 -p 9030:9030 apache/doris:all-in-one-4.1.3 +``` + +Wait for the frontend to become healthy, then connect with the MySQL protocol (port 9030, user `root`, no password in the all-in-one image). + +Create a test table with rows: + +```sql +CREATE DATABASE rustfs_demo; +CREATE TABLE rustfs_demo.events + (id INT, name VARCHAR(50)) + DISTRIBUTED BY HASH(id) BUCKETS 1 + PROPERTIES ("replication_num" = "1"); +INSERT INTO rustfs_demo.events VALUES (1, 'doris-on-rustfs'), (2, 'backup-test'); +``` + +Create the S3 repository, replacing the endpoint with the IP address of the RustFS container and both credential placeholders: + +```sql +CREATE REPOSITORY `rustfs_repo` + WITH S3 + ON LOCATION "s3://doris-backups/rustfs-repo" + PROPERTIES ( + "AWS_ENDPOINT" = "http://:9000", + "AWS_ACCESS_KEY" = "", + "AWS_SECRET_KEY" = "", + "AWS_REGION" = "us-east-1", + "AWS_PATH_STYLE_ACCESS" = "true" + ); +``` + +Doris 4.1 resolves the bucket into the endpoint hostname even with `AWS_PATH_STYLE_ACCESS` enabled, so a hostname endpoint fails with `UnknownHostException: doris-backups.rustfs`. Using the container IP address forces path-style requests and works; `SHOW REPOSITORIES` confirms the repository registered with an empty `ErrMsg`. + +## 2. Back up a table to RustFS + +Take a snapshot of the table: + +```sql +BACKUP SNAPSHOT rustfs_demo.demo_snapshot + TO rustfs_repo + ON (events); +``` + +The statement returns immediately; the backup job runs in the background. Watch its state: + +```sql +SHOW BACKUP; +``` + +Wait until `State` reaches `FINISHED` — the snapshot metadata and tablet data files are now objects in the bucket. + +## 3. Verify the backup in RustFS + +List the bucket: + +```bash +rc ls rustfs/doris-backups/ -r +``` + +The repository stores a repository descriptor, the snapshot metadata, and the tablet files: + +```text +rustfs-repo/__palo_repository_rustfs_repo/__repo_info +rustfs-repo/__palo_repository_rustfs_repo/__ss_demo_snapshot/__meta.d50ecf9b... +rustfs-repo/__palo_repository_rustfs_repo/__ss_demo_snapshot/__ss_content/.../...dat... +``` + +![Doris backup objects stored in the RustFS Console](./images/rustfs-doris-backup.png) + +## 4. Restore the table from RustFS + +Drop the table and restore it from the snapshot. The timestamp comes from the snapshot name shown by `SHOW SNAPSHOT ON REPOSITORY rustfs_repo;`: + +```sql +DROP TABLE rustfs_demo.events; + +RESTORE SNAPSHOT rustfs_demo.demo_snapshot + FROM rustfs_repo + ON (events) + PROPERTIES ( + "backup_timestamp" = "2026-09-21-16-35-12", + "replication_num" = "1" + ); +``` + +Wait for the restore job to finish and confirm the data: + +```sql +SHOW RESTORE; +SELECT count(*) FROM rustfs_demo.events; +``` + +```text +2 +``` + +## 5. Stop or reset the deployment + +Stop Doris while keeping the data: + +```bash +docker rm -f doris +``` + +The backup stays in the `doris-backups` bucket and can be restored into any Doris cluster that registers the same repository. To delete it, remove the bucket: + +```bash +rc rb rustfs/doris-backups --force +``` + +## Troubleshooting + +### `UnknownHostException: doris-backups.rustfs` when creating the repository + +Doris is building a virtual-hosted hostname from the bucket and endpoint. Use the RustFS container IP address in `AWS_ENDPOINT` together with `AWS_PATH_STYLE_ACCESS = "true"`, as shown above. + +### The backup stays in `SNAPSHOTING` for a long time + +The backend uploads the tablet files. Confirm the backend is healthy (`SHOW BACKENDS;`) and can reach the endpoint; the all-in-one image needs a minute or two after start before both processes report ready. + +### `Failed to create repository: ... file status` + +The bucket does not exist or the credentials are wrong. Create `doris-backups` with `rc mb` and re-check the access key pair. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Doris operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Doris backup and restore documentation](https://doris.apache.org/docs/data-operate/backup-restore/) to schedule periodic snapshots. diff --git a/content/ja/developer/integration/big-data/images/rustfs-clickhouse-disk.png b/content/ja/developer/integration/big-data/images/rustfs-clickhouse-disk.png new file mode 100644 index 00000000..9225bf66 Binary files /dev/null and b/content/ja/developer/integration/big-data/images/rustfs-clickhouse-disk.png differ diff --git a/content/ja/developer/integration/big-data/images/rustfs-doris-backup.png b/content/ja/developer/integration/big-data/images/rustfs-doris-backup.png new file mode 100644 index 00000000..9a936276 Binary files /dev/null and b/content/ja/developer/integration/big-data/images/rustfs-doris-backup.png differ diff --git a/content/ja/developer/integration/big-data/images/rustfs-opendal-objects.png b/content/ja/developer/integration/big-data/images/rustfs-opendal-objects.png new file mode 100644 index 00000000..d7129423 Binary files /dev/null and b/content/ja/developer/integration/big-data/images/rustfs-opendal-objects.png differ diff --git a/content/ja/developer/integration/big-data/images/rustfs-zeppelin-notebook.png b/content/ja/developer/integration/big-data/images/rustfs-zeppelin-notebook.png new file mode 100644 index 00000000..9c5fe782 Binary files /dev/null and b/content/ja/developer/integration/big-data/images/rustfs-zeppelin-notebook.png differ diff --git a/content/ja/developer/integration/big-data/index.md b/content/ja/developer/integration/big-data/index.md index a64033c9..cca193e0 100644 --- a/content/ja/developer/integration/big-data/index.md +++ b/content/ja/developer/integration/big-data/index.md @@ -7,11 +7,14 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo ## Systems +- [ClickHouse](./clickhouse.md) - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) - [Milvus](./milvus.md) - [MLflow](./mlflow.md) +- [OpenDAL](./opendal.md) - [DuckDB](./duckdb.md) +- [Doris](./doris.md) - [InfluxDB](./influxdb.md) - [Spark](./spark.md) - [Flink](./flink.md) diff --git a/content/ja/developer/integration/big-data/meta.json b/content/ja/developer/integration/big-data/meta.json index ebbb0f87..114ef68b 100644 --- a/content/ja/developer/integration/big-data/meta.json +++ b/content/ja/developer/integration/big-data/meta.json @@ -1,14 +1,18 @@ { "title": "データ分析", "pages": [ + "clickhouse", "iceberg", "pyiceberg", "milvus", "mlflow", + "opendal", "duckdb", + "doris", "influxdb", "spark", "flink", - "trino" + "trino", + "zeppelin" ] } diff --git a/content/ja/developer/integration/big-data/opendal.md b/content/ja/developer/integration/big-data/opendal.md new file mode 100644 index 00000000..074b7475 --- /dev/null +++ b/content/ja/developer/integration/big-data/opendal.md @@ -0,0 +1,108 @@ +--- +title: "OpenDAL" +description: "Access RustFS objects from applications through the Apache OpenDAL data access layer." +--- + +This guide connects [Apache OpenDAL](https://github.com/apache/opendal) — the unified data access layer — to **RustFS** through its `s3` service. You will run the OpenDAL Python binding against RustFS, write and read an object, list a prefix, and delete it. The workflow was verified with the `opendal` Python package 0.46 on `python:3.12-slim` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and Python 3.9 or later. OpenDAL supports the same `s3` service from Rust, Java, Node.js, and Go bindings with equivalent settings. + +## Architecture + +```mermaid +flowchart LR + App["Application"] -->|"Operator API"| OpenDAL["OpenDAL"] + OpenDAL -->|"s3 service"| RustFS["RustFS :9000"] +``` + +An OpenDAL `Operator` bound to the `s3` service exposes one uniform API — `write`, `read`, `stat`, `list`, `delete` — over the bucket, so the same code runs against S3, RustFS, or any other supported service by changing the connection settings. + +## 1. Set up the project + +Install the Python binding: + +```bash +pip install opendal +``` + +Create the script, replacing all connection placeholders: + +```python title="opendal_demo.py" +import opendal + +op = opendal.Operator( + "s3", + endpoint="http://:9000", + bucket="my-bucket", + access_key_id="", + secret_access_key="", + region="us-east-1", +) +op.write("opendal-demo/hello.txt", b"hello from opendal against rustfs") +print("read-back:", op.read("opendal-demo/hello.txt")) +print("content_length:", op.stat("opendal-demo/hello.txt").content_length) +for entry in op.list("opendal-demo/"): + print("listed:", entry.path) +op.delete("opendal-demo/hello.txt") +print("deleted:", not op.exists("opendal-demo/hello.txt")) +``` + +The credential option names are `access_key_id` and `secret_access_key` — the shorter `access_key` names do not exist and fail with a signing error. Path-style addressing is the default for non-AWS endpoints. + +## 2. Run the demo + +Run the script from a machine that can reach RustFS: + +```bash +python opendal_demo.py +``` + +```text +read-back: b"hello from opendal against rustfs" +content_length: 33 +listed: opendal-demo/hello.txt +deleted: True +``` + +The round trip exercises the full object lifecycle: write uploads the bytes, `read` fetches them back, `stat` returns the object size, `list` enumerates the prefix, and `delete` removes the object. + +## 3. Verify objects in RustFS + +Comment out the final `op.delete` line, run the script again, and list the prefix in RustFS: + +```bash +rc ls rustfs/my-bucket/opendal-demo/ -r +``` + +```text +hello.txt +data/rows.csv +``` + +![OpenDAL objects stored in the RustFS Console](./images/rustfs-opendal-objects.png) + +The objects visible in the RustFS Console are exactly the paths the OpenDAL API wrote. + +## 4. Stop or reset + +OpenDAL is a library and holds no state of its own. To clean up the demo objects: + +```bash +rc rm rustfs/my-bucket/opendal-demo/ --recursive --force +``` + +## Troubleshooting + +### `failed to load signing credential` + +The operator received no usable credentials. Use the exact option names `access_key_id` and `secret_access_key`; other spellings are silently ignored and signing then fails. + +### Connection or DNS errors on write + +Confirm the endpoint includes the scheme and port and is reachable from the application. Inside a Compose network the hostname is `rustfs`; from the host use `http://localhost:9000`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional OpenDAL operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [OpenDAL documentation](https://opendal.apache.org/docs/) to use the same operator from Rust, Java, or Node.js. diff --git a/content/ja/developer/integration/big-data/zeppelin.md b/content/ja/developer/integration/big-data/zeppelin.md new file mode 100644 index 00000000..e7f47375 --- /dev/null +++ b/content/ja/developer/integration/big-data/zeppelin.md @@ -0,0 +1,109 @@ +--- +title: "Apache Zeppelin" +description: "Store Apache Zeppelin notebooks in RustFS through the S3 notebook repository." +--- + +This guide connects [Apache Zeppelin](https://github.com/apache/zeppelin) — the web-based notebook for data analytics — to **RustFS** through Zeppelin's S3 notebook storage. You will start Zeppelin pointed at a RustFS bucket, create a note, and verify that the notebook file is stored in RustFS and survives a restart. The workflow was verified with `apache/zeppelin:0.12.0` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Browser["Browser"] -->|"notebook edits"| Z["Zeppelin :8080"] + Z -->|".zpln files"| RustFS["RustFS :9000"] +``` + +With the `S3NotebookRepo` storage class, every note is persisted as a `.zpln` JSON file under the `user/notebook/` prefix of the bucket. Zeppelin reads and writes the bucket directly, so notes survive container restarts and can be shared between instances. + +## 1. Start Zeppelin + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +AWS_ACCESS_KEY_ID= +AWS_SECRET_ACCESS_KEY= +``` + +Start Zeppelin on the same Docker network as RustFS with the S3 storage settings: + +```bash +docker run -d --name zeppelin --network oo-rustfs_default \ + -p 8080:8080 \ + -e AWS_ACCESS_KEY_ID \ + -e AWS_SECRET_ACCESS_KEY \ + -e ZEPPELIN_NOTEBOOK_STORAGE=org.apache.zeppelin.notebook.repo.S3NotebookRepo \ + -e ZEPPELIN_NOTEBOOK_S3_BUCKET=my-bucket \ + -e ZEPPELIN_NOTEBOOK_S3_ENDPOINT=http://rustfs:9000 \ + -e ZEPPELIN_NOTEBOOK_S3_PATH_STYLE_ACCESS=true \ + apache/zeppelin:0.12.0 +``` + +Zeppelin reads `ZEPPELIN_*` environment variables as configuration properties, so no `zeppelin-site.xml` edit is needed. This guide uses the existing `my-bucket`; notes land under its `user/notebook/` prefix, which the S3 storage creates on demand. + +## 2. Create a note + +Wait for the UI on `http://localhost:8080`, then create a note named `rustfs-demo` in the notebook list and add a paragraph, or use the REST API: + +```bash +NOTE=$(curl -s -X POST "http://localhost:8080/api/notebook" \ + -H "Content-Type: application/json" \ + -d '{"name": "rustfs-demo"}' | python3 -c "import json,sys; print(json.load(sys.stdin)['body'])") +echo "note id: $NOTE" +``` + +## 3. Verify the notebook in RustFS + +List the notebook prefix: + +```bash +rc ls rustfs/my-bucket/user/notebook/ -r +``` + +The note is stored as a JSON file named after the note and its ID: + +```text +user/notebook/rustfs-demo_2N4PY7UY5.zpln +``` + +![Zeppelin notebooks stored in the RustFS Console](./images/rustfs-zeppelin-notebook.png) + +Notes survive a restart because Zeppelin loads them from the bucket: + +```bash +docker restart zeppelin +curl -s "http://localhost:8080/api/notebook" | head -c 200 +``` + +The note list contains `2N4PY7UY5` again after the restart. + +## 4. Stop or reset the deployment + +Stop Zeppelin while keeping the notes: + +```bash +docker rm -f zeppelin +``` + +The notes stay in the `user/notebook/` prefix of `my-bucket`. To delete them, remove the prefix: + +```bash +rc rm rustfs/my-bucket/user/notebook/ --recursive --force +``` + +## Troubleshooting + +### Zeppelin starts but notes never appear in the bucket + +Confirm the three `ZEPPELIN_NOTEBOOK_S3_*` variables are set and that the credentials environment variables reach the container — the S3 repository is initialized at startup, so the container must be recreated after any change. + +### `UnknownHostException: my-bucket.rustfs` + +The path-style flag was not picked up. Keep `ZEPPELIN_NOTEBOOK_S3_PATH_STYLE_ACCESS=true` exactly as shown; with path-style disabled Zeppelin treats the bucket as a hostname. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Zeppelin operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Zeppelin storage documentation](https://zeppelin.apache.org/docs/latest/setup/storage/storage.html#notebook-storage-in-s3) to organize notes in per-user prefixes. diff --git a/content/ja/developer/integration/devops/images/rustfs-jenkins-artifacts.png b/content/ja/developer/integration/devops/images/rustfs-jenkins-artifacts.png new file mode 100644 index 00000000..19c8cf87 Binary files /dev/null and b/content/ja/developer/integration/devops/images/rustfs-jenkins-artifacts.png differ diff --git a/content/ja/developer/integration/devops/index.md b/content/ja/developer/integration/devops/index.md index caae08a1..9e31d321 100644 --- a/content/ja/developer/integration/devops/index.md +++ b/content/ja/developer/integration/devops/index.md @@ -9,6 +9,7 @@ S3 互換エンドポイントをサポートする DevOps プラットフォー - [Elasticsearch](./elasticsearch.md) - [Gitea](./gitea.md) +- [Jenkins](./jenkins.md) - [Terraform](./terraform.md) アーティファクト、ステート、テレメトリデータは専用バケットに保存し、必要なバケット操作のみに権限が絞られた認証情報を使用してください。 diff --git a/content/ja/developer/integration/devops/jenkins.md b/content/ja/developer/integration/devops/jenkins.md new file mode 100644 index 00000000..1d78f31b --- /dev/null +++ b/content/ja/developer/integration/devops/jenkins.md @@ -0,0 +1,144 @@ +--- +title: "Jenkins" +description: "Store Jenkins build artifacts in RustFS with the Artifact Manager on S3 plugin." +--- + +This guide connects [Jenkins](https://github.com/jenkinsci/jenkins) — the automation server — to **RustFS** through the Artifact Manager on S3 plugin. You will start Jenkins with the plugin, point its artifact manager at a RustFS bucket, run a job that archives an artifact, and verify that the artifact is stored in RustFS. The workflow was verified with `jenkins/jenkins:lts-jdk17` (Jenkins 2.5xx LTS) and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Dev["Developer"] -->|"trigger build"| J["Jenkins :8080"] + J -->|"archiveArtifacts"| RustFS["RustFS :9000"] +``` + +With the plugin active, every job that publishes artifacts through the standard `archiveArtifacts` step (or `stash`/`unstash`) uploads them to the `jenkins-artifacts` bucket under the configured prefix instead of storing them on the controller disk. + +## 1. Create the project files + +Create the bucket first — the plugin validates the bucket but does not create it: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/jenkins-artifacts +``` + +Create a Dockerfile that installs the plugin into the Jenkins LTS image: + +```dockerfile title="Dockerfile" +FROM jenkins/jenkins:lts-jdk17 +USER root +RUN jenkins-plugin-cli --plugins artifact-manager-s3 aws-credentials +USER jenkins +``` + +The `artifact-manager-s3` plugin brings in the AWS credentials support it needs; listing `aws-credentials` explicitly keeps the credential type available. + +Build and start Jenkins on the same Docker network as RustFS: + +```bash +docker build -t jenkins-rustfs . +docker run -d --name jenkins --network oo-rustfs_default \ + -p 8080:8080 -v jenkins-home:/var/jenkins_home jenkins-rustfs +``` + +Complete the setup wizard, then create an AWS credential: **Manage Jenkins → Credentials → global → Add Credentials**, kind **AWS Credential**, ID `rustfs-creds`, with your RustFS access key and secret key. + +## 2. Configure the artifact manager + +Open **Manage Jenkins → AWS Configuration** (from the `aws-global-configuration` plugin) and set: + +- **Region name**: `us-east-1` +- **Credentials**: `rustfs-creds` + +Open **Manage Jenkins → System** and locate the **Artifact Management for Builds** section. Select **Cloud Provider Amazon S3**, and fill in the S3 configuration: + +- **S3 Bucket Name**: `jenkins-artifacts` +- **S3 Bucket Region**: `us-east-1` +- **Base Prefix**: `artifacts/` +- **Custom Endpoint**: `:9000` (for example `rustfs:9000` inside the Compose network) +- **Custom Signing Region**: `us-east-1` +- **Use Path Style URL**: enabled +- **Use Insecure HTTP**: enabled +- **Disable Session Token**: enabled + +Path-style addressing and plain HTTP are required for a non-AWS endpoint without TLS. Disabling the session token stops the plugin from calling AWS STS, which a plain access key pair cannot answer. Click **Validate S3 Bucket configuration** to confirm the settings, then save. + +## 3. Run a job that archives an artifact + +Create a freestyle job (or pipeline) that produces a file and archives it: + +```groovy +pipeline { + agent any + stages { + stage('Build') { + steps { + sh 'echo "jenkins artifact stored on rustfs" > report.txt' + } + } + } + post { + always { + archiveArtifacts 'report.txt' + } + } +} +``` + +Run the build and wait for it to finish. The artifact upload goes to RustFS transparently — the job configuration does not mention S3 at all. + +## 4. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/jenkins-artifacts/ -r +``` + +The artifact is stored under the prefix, organized by job and build number: + +```text +artifacts/s3-artifacts-demo/3/artifacts/report.txt +``` + +![Jenkins artifacts stored in the RustFS Console](./images/rustfs-jenkins-artifacts.png) + +Downloading the artifact from the build page reads it back from RustFS. + +## 5. Stop or reset the deployment + +Stop Jenkins while keeping the data: + +```bash +docker rm -f jenkins +``` + +The artifacts stay in the `jenkins-artifacts` bucket. To delete them, remove the bucket: + +```bash +rc rb rustfs/jenkins-artifacts --force +``` + +## Troubleshooting + +### `StsException: The security token included in the request is invalid` + +The plugin is calling AWS STS to obtain session credentials. Enable **Disable Session Token** in the S3 configuration — a static access key pair cannot answer an STS call. + +### `UnknownHostException: jenkins-artifacts.rustfs` + +The plugin is using virtual-hosted addressing. Enable **Use Path Style URL** — RustFS resolves buckets from the URL path, not the hostname. + +### `No valid session credentials` or empty credential errors + +Confirm the AWS Configuration page has the credential selected and saved **before** the S3 bucket settings are used, and that the credential ID matches the one you created. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Jenkins integrations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Artifact Manager on S3 plugin documentation](https://plugins.jenkins.io/artifact-manager-s3/) for stash support and cleanup options. diff --git a/content/ja/developer/integration/devops/meta.json b/content/ja/developer/integration/devops/meta.json index 61ad9fb8..67c84c6e 100644 --- a/content/ja/developer/integration/devops/meta.json +++ b/content/ja/developer/integration/devops/meta.json @@ -3,6 +3,7 @@ "pages": [ "elasticsearch", "gitea", + "jenkins", "terraform" ] } diff --git a/content/ja/developer/integration/index.md b/content/ja/developer/integration/index.md index e13a8768..5e47330c 100644 --- a/content/ja/developer/integration/index.md +++ b/content/ja/developer/integration/index.md @@ -9,10 +9,10 @@ description: "RustFS をリバースプロキシ、バックアップツール - [Reverse Proxy](./reverse-proxy/index.md) は Nginx、Traefik、Caddy、HAProxy を扱います。 - [Backup](./backup/index.md) は Restic と Longhorn を扱います。 -- [データ分析](./big-data/index.md) は Iceberg を扱います。 -- [オブザーバビリティ](./observability/index.md) は OpenObserve を扱います。 +- [データ分析](./big-data/index.md) は ClickHouse、Doris、Iceberg、Milvus、OpenDAL、Zeppelin などの分析システムを扱います。 +- [オブザーバビリティ](./observability/index.md) は Fluentd、OpenObserve、OpenTelemetry、Thanos、Tempo などのテレメトリシステムを扱います。 - [その他](./others/index.md) はコミュニティ主導の Python 用 capo SDK を扱います。 - [コンテナレジストリ](./registry/index.md) は Harbor を扱います。 -- [DevOps](./devops/index.md) は Elasticsearch、Gitea、Terraform を扱います。 +- [DevOps](./devops/index.md) は Elasticsearch、Gitea、Jenkins、Terraform を扱います。 各ガイドでは、連携先システムを設定する際に使用する RustFS のエンドポイントとアドレス指定の要件を示します。 \ No newline at end of file diff --git a/content/ja/developer/integration/observability/fluentd.md b/content/ja/developer/integration/observability/fluentd.md new file mode 100644 index 00000000..4607c0ec --- /dev/null +++ b/content/ja/developer/integration/observability/fluentd.md @@ -0,0 +1,152 @@ +--- +title: "Fluentd" +description: "Ship Fluentd log events to RustFS with the S3 output plugin." +--- + +This guide connects [Fluentd](https://github.com/fluent/fluentd) — the open-source data collector — to **RustFS** through the `out_s3` output plugin. You will run Fluentd with a tail source, buffer log events, and verify that the flushed objects are stored in RustFS. The workflow was verified with `fluent/fluentd:v1.17-1`, `fluent-plugin-s3` 1.8.6, and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + App["Application"] -->|"writes lines"| File["app.log"] + File -->|tail| Fluentd["Fluentd"] + Fluentd -->|"gzip objects"| RustFS["RustFS :9000"] +``` + +The tail source reads new lines from `app.log` and hands them to the S3 output, which buffers events on a time key and uploads a gzip object per flush window. + +## 1. Create the project files + +Create the bucket first — Fluentd does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/fluentd-data +``` + +The official Fluentd image runs as a non-root user and cannot install gems at startup, so build a small image with the S3 plugin: + +```dockerfile title="Dockerfile" +FROM fluent/fluentd:v1.17-1 +USER root +RUN gem install fluent-plugin-s3 --no-document +USER fluent +``` + +Create the Fluentd configuration, replacing both credential placeholders: + +```nginx title="fluent.conf" + + @type tail + path /var/log/app.log + pos_file /var/log/app.log.pos + tag rustfs.demo + + @type none + + + + + @type s3 + aws_key_id + aws_sec_key + s3_bucket fluentd-data + s3_endpoint http://rustfs:9000/ + s3_region us-east-1 + force_path_style true + path fluentd-logs + + @type memory + timekey 30s + timekey_wait 0s + flush_mode immediate + + +``` + +`force_path_style true` is required — without it the plugin constructs `fluentd-data.rustfs` as a hostname and every request fails with a DNS error. Inside the Compose network the hostname is `rustfs`; from the host use `http://localhost:9000/`. + +Build the image and start Fluentd on the same Docker network as RustFS: + +```bash +docker build -t fluentd-rustfs . +mkdir -p logs +docker run -d --name fluentd --network oo-rustfs_default \ + -v "$PWD/fluent.conf":/fluentd/etc/fluent.conf:ro \ + -v "$PWD/logs":/var/log fluentd-rustfs +``` + +## 2. Produce log events + +Append lines to the watched file: + +```bash +echo "rustfs fluentd demo line 1" >> logs/app.log +echo "rustfs fluentd demo line 2" >> logs/app.log +``` + +With a 30-second time key and immediate flush mode, each window uploads one gzip object shortly after it closes. Wait about a minute. + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/fluentd-data/ -r +``` + +Each flush window produces one gzipped object: + +```text +fluentd-logs20260921154400_0.gz +fluentd-logs20260921154400_1.gz +``` + +Read one object back to confirm the events are intact: + +```bash +rc cat rustfs/fluentd-data/fluentd-logs20260921154400_0.gz | gunzip +``` + +![Fluentd log objects stored in the RustFS Console](./images/rustfs-fluentd-logs.png) + +## 4. Stop or reset the deployment + +Stop Fluentd while keeping the data: + +```bash +docker rm -f fluentd +``` + +The objects stay in the `fluentd-data` bucket. To delete them, remove the bucket: + +```bash +rc rb rustfs/fluentd-data --force +``` + +## Troubleshooting + +### `Unknown output plugin 's3'` + +The plugin is not installed. Confirm the Dockerfile installs `fluent-plugin-s3` as `root` before switching back to the `fluent` user — installing at container start as the default user fails with permission errors. + +### `Failed to open TCP connection to fluentd-data.rustfs` + +Virtual-hosted addressing is in use. Add `force_path_style true` to the `s3` output so the bucket stays in the URL path. + +### The worker crashes in a restart loop + +The output fails hard when the bucket does not exist. Create `fluentd-data` before starting Fluentd and check the startup log: + +```bash +docker logs fluentd | grep -iE "error|bucket" | tail +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Fluentd outputs. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [fluent-plugin-s3 documentation](https://github.com/fluent/fluent-plugin-s3) for object key formats and compression options. diff --git a/content/ja/developer/integration/observability/images/rustfs-fluentd-logs.png b/content/ja/developer/integration/observability/images/rustfs-fluentd-logs.png new file mode 100644 index 00000000..3ab768ca Binary files /dev/null and b/content/ja/developer/integration/observability/images/rustfs-fluentd-logs.png differ diff --git a/content/ja/developer/integration/observability/images/rustfs-otel-logs.png b/content/ja/developer/integration/observability/images/rustfs-otel-logs.png new file mode 100644 index 00000000..2f68a22a Binary files /dev/null and b/content/ja/developer/integration/observability/images/rustfs-otel-logs.png differ diff --git a/content/ja/developer/integration/observability/index.md b/content/ja/developer/integration/observability/index.md index f6485166..56acb094 100644 --- a/content/ja/developer/integration/observability/index.md +++ b/content/ja/developer/integration/observability/index.md @@ -7,7 +7,9 @@ S3 互換エンドポイントをサポートするオブザーバビリティ ## プラットフォーム +- [Fluentd](./fluentd.md) - [OpenObserve](./openobserve.md) +- [OpenTelemetry](./opentelemetry.md) - [Loki](./loki.md) - [Tempo](./tempo.md) - [Thanos](./thanos.md) diff --git a/content/ja/developer/integration/observability/meta.json b/content/ja/developer/integration/observability/meta.json index 16c15ac9..1f09554f 100644 --- a/content/ja/developer/integration/observability/meta.json +++ b/content/ja/developer/integration/observability/meta.json @@ -1,8 +1,10 @@ { "title": "オブザーバビリティ", "pages": [ - "openobserve", + "fluentd", "loki", + "openobserve", + "opentelemetry", "tempo", "thanos" ] diff --git a/content/ja/developer/integration/observability/opentelemetry.md b/content/ja/developer/integration/observability/opentelemetry.md new file mode 100644 index 00000000..283df413 --- /dev/null +++ b/content/ja/developer/integration/observability/opentelemetry.md @@ -0,0 +1,145 @@ +--- +title: "OpenTelemetry" +description: "Export OpenTelemetry Collector logs to RustFS with the AWS S3 exporter." +--- + +This guide connects the [OpenTelemetry Collector](https://opentelemetry.io/docs/collector/) — the CNCF telemetry pipeline — to **RustFS** through the collector's `awss3` exporter. You will run the contrib collector with a `filelog` receiver, ship the log lines of a file into the `otel-data` bucket, and verify the partitioned objects in RustFS. The workflow was verified with `otel/opentelemetry-collector-contrib:0.138.0` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Log["Application log file"] -->|filelog| Collector["OTel Collector"] + Collector -->|awss3| RustFS["RustFS :9000"] +``` + +The `filelog` receiver tails the log file and the `awss3` exporter uploads batches to the bucket, partitioned into time-based prefixes. Metrics and traces can be routed through the same exporter with their own pipelines. + +## 1. Create the project files + +Create the bucket first — the exporter does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/otel-data +``` + +Create the collector configuration, replacing both credential placeholders: + +```yaml title="config.yaml" +receivers: + filelog: + include: [/var/log/app.log] + start_at: beginning + +exporters: + awss3: + s3uploader: + region: us-east-1 + s3_bucket: otel-data + endpoint: http://rustfs:9000 + s3_force_path_style: true + disable_ssl: true + file_prefix: logs/app + marshaler: body + +service: + pipelines: + logs: + receivers: [filelog] + exporters: [awss3] +``` + +The exporter reads credentials from the standard `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, and `AWS_REGION` environment variables. `marshaler: body` writes each log line as plain text; omit it to store OTLP JSON instead. `s3_force_path_style` and `disable_ssl` are required for a plain-HTTP, non-AWS endpoint. + +Create the Compose file: + +```yaml title="compose.yaml" +services: + collector: + image: otel/opentelemetry-collector-contrib:0.138.0 + command: ["--config=/etc/otelcol-contrib/config.yaml"] + environment: + AWS_ACCESS_KEY_ID: + AWS_SECRET_ACCESS_KEY: + AWS_REGION: us-east-1 + volumes: + - ./config.yaml:/etc/otelcol-contrib/config.yaml:ro + - ./app.log:/var/log/app.log + networks: + - otel + +networks: + otel: +``` + +## 2. Start the collector and produce logs + +Start the stack and append lines to the watched file: + +```bash +echo "otel demo log line one" > app.log +docker compose up -d +echo "second line after start" >> app.log +``` + +The collector tails the file from the beginning and uploads each buffer when the partition rolls over, so allow about a minute after the last line before checking. + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/otel-data/ -r +``` + +Log records are uploaded under time-based partitions with the configured file prefix: + +```text +year=2026/month=09/day=21/hour=15/minute=53/logs/applogs_288361608.txt +``` + +Read one object back to confirm the lines arrived intact: + +```bash +rc cat rustfs/otel-data/year=2026/month=09/day=21/hour=15/minute=53/logs/applogs_288361608.txt +``` + +```text +otel demo log line one +second line after start +``` + +![OpenTelemetry log objects stored in the RustFS Console](./images/rustfs-otel-logs.png) + +## 4. Stop or reset the deployment + +Stop the collector while keeping the data: + +```bash +docker compose down +``` + +The objects stay in the `otel-data` bucket. To delete them, remove the bucket: + +```bash +rc rb rustfs/otel-data --force +``` + +## Troubleshooting + +### `has invalid keys` when the collector starts + +The `awss3` exporter schema differs between collector releases. Version 0.138 nests the upload settings under `s3uploader` as shown above; the endpoint key is `endpoint` (not `s3_endpoint`). Newer releases move these keys to the top level — check the README for your exact collector version. + +### Nothing appears in the bucket + +Confirm the credentials environment variables are set on the collector container, that `s3_force_path_style` is `true`, and that the exporter can reach `http://rustfs:9000` from inside the Compose network. Enable the `debug` exporter on the same pipeline to see whether records flow at all. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional collector operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [AWS S3 exporter documentation](https://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/exporter/awss3exporter) to add metrics and traces pipelines. diff --git a/content/zh/developer/integration/big-data/clickhouse.md b/content/zh/developer/integration/big-data/clickhouse.md new file mode 100644 index 00000000..00f8f12a --- /dev/null +++ b/content/zh/developer/integration/big-data/clickhouse.md @@ -0,0 +1,160 @@ +--- +title: "ClickHouse" +description: "运行 ClickHouse,以 RustFS 作为 MergeTree 表数据的 S3 磁盘。" +--- + +本指南通过 ClickHouse 的 S3 磁盘存储策略,将实时 OLAP 数据库 [ClickHouse](https://github.com/ClickHouse/ClickHouse) 连接到 **RustFS**。你将使用 Docker 启动 ClickHouse 服务器,创建一张数据部分存储在 RustFS 上的 MergeTree 表,插入数据行,并验证表数据已存入存储桶。整个流程使用 `clickhouse/clickhouse-server:25.8` 和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Client["SQL client"] -->|"queries"| CH["ClickHouse :8123"] + CH -->|"MergeTree parts"| RustFS["RustFS :9000"] +``` + +`rustfs` 磁盘是指向 `clickhouse-data` 存储桶的 ClickHouse S3 磁盘。使用对应存储策略创建的表会把数据部分——数据、索引和校验文件——写入存储桶而非本地文件系统。 + +## 1. 创建项目文件 + +先创建存储桶——ClickHouse 不会创建桶: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/clickhouse-data +``` + +创建存储配置,并替换两个凭证占位符: + +```xml title="storage.xml" + + + + + s3 + http://rustfs:9000/clickhouse-data/ + + + + + + + +
+ rustfs +
+
+
+
+
+
+``` + +端点必须以 `/` 结尾,且桶名作为路径的第一段。Compose 网络内主机名为 `rustfs`;宿主机上使用 (`http://localhost:9000/clickhouse-data/`)。 + +挂载配置启动 ClickHouse: + +```bash +docker run -d --name clickhouse --network oo-rustfs_default \ + -p 8123:8123 \ + -e CLICKHOUSE_PASSWORD= \ + -v "$PWD/storage.xml":/etc/clickhouse-server/config.d/storage.xml:ro \ + clickhouse/clickhouse-server:25.8 +``` + +## 2. 在 S3 磁盘上创建表 + +等待 HTTP 接口就绪,然后创建数据库和带存储策略的 MergeTree 表: + +```bash +curl "http://localhost:8123/?password=" \ + --data-binary "CREATE DATABASE rustfs_demo" + +curl "http://localhost:8123/?password=" \ + --data-binary "CREATE TABLE rustfs_demo.events + (id UInt32, name String) + ENGINE = MergeTree ORDER BY id + SETTINGS storage_policy = 'rustfs_policy'" + +curl "http://localhost:8123/?password=" \ + --data-binary "INSERT INTO rustfs_demo.events + VALUES (1, 'clickhouse-on-rustfs'), (2, 'second')" +``` + +读回数据行,并确认 ClickHouse 报告数据部分位于 `rustfs` 磁盘上: + +```bash +curl "http://localhost:8123/?password=" \ + --data-binary "SELECT count(), any(name) FROM rustfs_demo.events" + +curl "http://localhost:8123/?password=" \ + --data-binary "SELECT name, disk_name FROM system.parts + WHERE database = 'rustfs_demo' AND active" +``` + +```text +2 clickhouse-on-rustfs +all_1_1_0 rustfs +``` + +## 3. 在 RustFS 中验证对象 + +列出存储桶: + +```bash +rc ls rustfs/clickhouse-data/ -r +``` + +ClickHouse 把每个数据部分写成内容寻址的 blob。输出包含若干小对象,写入的部分越多对象越多: + +```text +dtg/hpsyncexixvdnsgorvseobogcgowg +dzp/zfblobhsatzdveqdrsfqcupkdehja +izg/gvhchqobrizpkdftuvlmakkoutwps +``` + +![RustFS 控制台中存储的 ClickHouse 数据部分](./images/rustfs-clickhouse-disk.png) + +数据部分存放在 RustFS 中,因此容器重启后数据依然可查: + +```bash +docker restart clickhouse +curl "http://localhost:8123/?password=" \ + --data-binary "SELECT count() FROM rustfs_demo.events" +``` + +## 4. 停止或重置部署 + +停止服务器并保留数据: + +```bash +docker rm -f clickhouse +``` + +数据部分保留在 `clickhouse-data` 存储桶中,下次启动后即可继续查询。若要删除数据,请移除存储桶: + +```bash +rc rb rustfs/clickhouse-data --force +``` + +## 故障排查 + +### 每次查询都返回 `REQUIRED_PASSWORD` + +ClickHouse 25.8 镜像要求 `default` 用户使用密码。在容器上设置 `CLICKHOUSE_PASSWORD`,并在查询时传入相同的 `password` 参数,如上所示。 + +### 建表时报磁盘或端点错误 + +确认建表前存储桶已存在、端点以 `/` 结尾、凭证与 RustFS 部署一致。查看服务器日志中的底层 S3 错误: + +```bash +docker logs clickhouse | grep -i s3 | tail +``` + +## 后续步骤 + +- 在采用更多 ClickHouse 操作之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [ClickHouse S3 磁盘文档](https://clickhouse.com/docs/engines/table-engines/mergetree-family/mergetree#table_engine-mergetree-s3)添加缓存磁盘或冷热分层策略。 diff --git a/content/zh/developer/integration/big-data/doris.md b/content/zh/developer/integration/big-data/doris.md new file mode 100644 index 00000000..bc58d1a7 --- /dev/null +++ b/content/zh/developer/integration/big-data/doris.md @@ -0,0 +1,161 @@ +--- +title: "Apache Doris" +description: "通过 S3 仓库把 Apache Doris 表备份到 RustFS 并恢复。" +--- + +本指南通过 S3 备份仓库,将实时分析型数据仓库 [Apache Doris](https://github.com/apache/doris) 连接到 **RustFS**。你将启动一个 all-in-one Doris 容器,创建指向 RustFS 存储桶的 S3 仓库,备份一张表,删除后再从 RustFS 恢复。整个流程使用 `apache/doris:all-in-one-4.1.3` 和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Client["SQL client"] -->|"queries"| Doris["Doris FE/BE"] + Doris -->|"BACKUP / RESTORE"| RustFS["RustFS :9000"] +``` + +仓库是 `doris-backups` 存储桶下的一个命名 S3 位置。`BACKUP SNAPSHOT` 上传表元数据和 tablet 数据文件;`RESTORE SNAPSHOT` 把它们下载为新表。 + +## 1. 启动 Doris 并创建仓库 + +先创建存储桶——Doris 不会创建桶: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/doris-backups +``` + +在与 RustFS 相同的 Docker 网络中启动 all-in-one 容器: + +```bash +docker run -d --name doris --network oo-rustfs_default \ + -p 8030:8030 -p 9030:9030 apache/doris:all-in-one-4.1.3 +``` + +等待前端健康后,用 MySQL 协议连接(端口 9030,用户 `root`,all-in-one 镜像无密码)。 + +创建带数据的测试表: + +```sql +CREATE DATABASE rustfs_demo; +CREATE TABLE rustfs_demo.events + (id INT, name VARCHAR(50)) + DISTRIBUTED BY HASH(id) BUCKETS 1 + PROPERTIES ("replication_num" = "1"); +INSERT INTO rustfs_demo.events VALUES (1, 'doris-on-rustfs'), (2, 'backup-test'); +``` + +创建 S3 仓库,把端点替换为 RustFS 容器的 IP 地址,并替换两个凭证占位符: + +```sql +CREATE REPOSITORY `rustfs_repo` + WITH S3 + ON LOCATION "s3://doris-backups/rustfs-repo" + PROPERTIES ( + "AWS_ENDPOINT" = "http://:9000", + "AWS_ACCESS_KEY" = "", + "AWS_SECRET_KEY" = "", + "AWS_REGION" = "us-east-1", + "AWS_PATH_STYLE_ACCESS" = "true" + ); +``` + +Doris 4.1 即使启用了 `AWS_PATH_STYLE_ACCESS` 也会把桶解析进端点主机名,因此主机名端点会报 `UnknownHostException: doris-backups.rustfs`。改用容器 IP 地址即可强制 path-style 请求;`SHOW REPOSITORIES` 确认仓库注册成功且 `ErrMsg` 为空。 + +## 2. 把表备份到 RustFS + +对表做快照: + +```sql +BACKUP SNAPSHOT rustfs_demo.demo_snapshot + TO rustfs_repo + ON (events); +``` + +该语句立即返回;备份作业在后台运行。观察其状态: + +```sql +SHOW BACKUP; +``` + +等待 `State` 到达 `FINISHED`——快照元数据和 tablet 数据文件此时已成为存储桶中的对象。 + +## 3. 在 RustFS 中验证备份 + +列出存储桶: + +```bash +rc ls rustfs/doris-backups/ -r +``` + +仓库中保存了仓库描述文件、快照元数据和 tablet 文件: + +```text +rustfs-repo/__palo_repository_rustfs_repo/__repo_info +rustfs-repo/__palo_repository_rustfs_repo/__ss_demo_snapshot/__meta.d50ecf9b... +rustfs-repo/__palo_repository_rustfs_repo/__ss_demo_snapshot/__ss_content/.../...dat... +``` + +![RustFS 控制台中存储的 Doris 备份对象](./images/rustfs-doris-backup.png) + +## 4. 从 RustFS 恢复表 + +删除表并从快照恢复。时间戳来自 `SHOW SNAPSHOT ON REPOSITORY rustfs_repo;` 显示的快照名: + +```sql +DROP TABLE rustfs_demo.events; + +RESTORE SNAPSHOT rustfs_demo.demo_snapshot + FROM rustfs_repo + ON (events) + PROPERTIES ( + "backup_timestamp" = "2026-09-21-16-35-12", + "replication_num" = "1" + ); +``` + +等待恢复作业完成并确认数据: + +```sql +SHOW RESTORE; +SELECT count(*) FROM rustfs_demo.events; +``` + +```text +2 +``` + +## 5. 停止或重置部署 + +停止 Doris 并保留数据: + +```bash +docker rm -f doris +``` + +备份保留在 `doris-backups` 存储桶中,任何注册了相同仓库的 Doris 集群都可以恢复它。若要删除备份,请移除存储桶: + +```bash +rc rb rustfs/doris-backups --force +``` + +## 故障排查 + +### 创建仓库时报 `UnknownHostException: doris-backups.rustfs` + +Doris 正在用桶和端点构造 virtual-hosted 主机名。按上文所示,在 `AWS_ENDPOINT` 中使用 RustFS 容器的 IP 地址并设置 `AWS_PATH_STYLE_ACCESS = "true"`。 + +### 备份长时间停留在 `SNAPSHOTING` + +后端正在上传 tablet 文件。确认后端健康(`SHOW BACKENDS;`)且能访问端点;all-in-one 镜像启动后需要一两分钟两个进程才全部就绪。 + +### `Failed to create repository: ... file status` + +存储桶不存在或凭证错误。用 `rc mb` 创建 `doris-backups`,并重新检查访问密钥对。 + +## 后续步骤 + +- 在采用更多 Doris 操作之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Doris 备份恢复文档](https://doris.apache.org/docs/data-operate/backup-restore/)配置周期性快照。 diff --git a/content/zh/developer/integration/big-data/images/rustfs-clickhouse-disk.png b/content/zh/developer/integration/big-data/images/rustfs-clickhouse-disk.png new file mode 100644 index 00000000..7ebbd441 Binary files /dev/null and b/content/zh/developer/integration/big-data/images/rustfs-clickhouse-disk.png differ diff --git a/content/zh/developer/integration/big-data/images/rustfs-doris-backup.png b/content/zh/developer/integration/big-data/images/rustfs-doris-backup.png new file mode 100644 index 00000000..8cdc2da6 Binary files /dev/null and b/content/zh/developer/integration/big-data/images/rustfs-doris-backup.png differ diff --git a/content/zh/developer/integration/big-data/images/rustfs-opendal-objects.png b/content/zh/developer/integration/big-data/images/rustfs-opendal-objects.png new file mode 100644 index 00000000..5564d4fb Binary files /dev/null and b/content/zh/developer/integration/big-data/images/rustfs-opendal-objects.png differ diff --git a/content/zh/developer/integration/big-data/images/rustfs-zeppelin-notebook.png b/content/zh/developer/integration/big-data/images/rustfs-zeppelin-notebook.png new file mode 100644 index 00000000..795821cd Binary files /dev/null and b/content/zh/developer/integration/big-data/images/rustfs-zeppelin-notebook.png differ diff --git a/content/zh/developer/integration/big-data/index.md b/content/zh/developer/integration/big-data/index.md index 73189d80..ecce2ee7 100644 --- a/content/zh/developer/integration/big-data/index.md +++ b/content/zh/developer/integration/big-data/index.md @@ -7,11 +7,14 @@ description: "通过 S3 兼容的对象存储接口将数据分析系统连接 ## 系统 +- [ClickHouse](./clickhouse.md) - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) - [Milvus](./milvus.md) - [MLflow](./mlflow.md) +- [OpenDAL](./opendal.md) - [DuckDB](./duckdb.md) +- [Doris](./doris.md) - [InfluxDB](./influxdb.md) - [Spark](./spark.md) - [Flink](./flink.md) diff --git a/content/zh/developer/integration/big-data/meta.json b/content/zh/developer/integration/big-data/meta.json index a951cb56..ecf05567 100644 --- a/content/zh/developer/integration/big-data/meta.json +++ b/content/zh/developer/integration/big-data/meta.json @@ -1,14 +1,18 @@ { "title": "数据分析", "pages": [ + "clickhouse", "iceberg", "pyiceberg", "milvus", "mlflow", + "opendal", "duckdb", + "doris", "influxdb", "spark", "flink", - "trino" + "trino", + "zeppelin" ] } diff --git a/content/zh/developer/integration/big-data/opendal.md b/content/zh/developer/integration/big-data/opendal.md new file mode 100644 index 00000000..d2a06de2 --- /dev/null +++ b/content/zh/developer/integration/big-data/opendal.md @@ -0,0 +1,108 @@ +--- +title: "OpenDAL" +description: "应用通过 Apache OpenDAL 数据访问层读写 RustFS 对象。" +--- + +本指南通过 `s3` 服务,将统一数据访问层 [Apache OpenDAL](https://github.com/apache/opendal) 连接到 **RustFS**。你将针对 RustFS 运行 OpenDAL 的 Python 绑定,写入并读回一个对象,列出前缀,再删除它。整个流程使用 `opendal` Python 包 0.46(运行于 `python:3.12-slim`)和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker 和 Python 3.9 或更高版本。OpenDAL 的 Rust、Java、Node.js 和 Go 绑定提供等效设置的相同 `s3` 服务。 + +## 架构 + +```mermaid +flowchart LR + App["Application"] -->|"Operator API"| OpenDAL["OpenDAL"] + OpenDAL -->|"s3 service"| RustFS["RustFS :9000"] +``` + +绑定到 `s3` 服务的 OpenDAL `Operator` 在存储桶之上提供统一 API——`write`、`read`、`stat`、`list`、`delete`——换一套连接设置即可在 S3、RustFS 或任何受支持的服务之间切换。 + +## 1. 准备项目 + +安装 Python 绑定: + +```bash +pip install opendal +``` + +创建脚本,并替换全部连接占位符: + +```python title="opendal_demo.py" +import opendal + +op = opendal.Operator( + "s3", + endpoint="http://:9000", + bucket="my-bucket", + access_key_id="", + secret_access_key="", + region="us-east-1", +) +op.write("opendal-demo/hello.txt", b"hello from opendal against rustfs") +print("read-back:", op.read("opendal-demo/hello.txt")) +print("content_length:", op.stat("opendal-demo/hello.txt").content_length) +for entry in op.list("opendal-demo/"): + print("listed:", entry.path) +op.delete("opendal-demo/hello.txt") +print("deleted:", not op.exists("opendal-demo/hello.txt")) +``` + +凭证参数名是 `access_key_id` 和 `secret_access_key`——较短的 `access_key` 写法不存在,会导致签名错误。对非 AWS 端点默认使用 path-style 寻址。 + +## 2. 运行演示 + +在能访问 RustFS 的机器上运行脚本: + +```bash +python opendal_demo.py +``` + +```text +read-back: b"hello from opendal against rustfs" +content_length: 33 +listed: opendal-demo/hello.txt +deleted: True +``` + +这次往返覆盖了完整的对象生命周期:`write` 上传字节,`read` 读回,`stat` 返回对象大小,`list` 枚举前缀,`delete` 删除对象。 + +## 3. 在 RustFS 中验证对象 + +注释掉最后的 `op.delete` 行,再次运行脚本,然后列出 RustFS 中的前缀: + +```bash +rc ls rustfs/my-bucket/opendal-demo/ -r +``` + +```text +hello.txt +data/rows.csv +``` + +![RustFS 控制台中存储的 OpenDAL 对象](./images/rustfs-opendal-objects.png) + +RustFS 控制台中可见的对象正是 OpenDAL API 写入的路径。 + +## 4. 停止或重置 + +OpenDAL 是一个类库,自身不保存状态。清理演示对象: + +```bash +rc rm rustfs/my-bucket/opendal-demo/ --recursive --force +``` + +## 故障排查 + +### `failed to load signing credential` + +Operator 没有获得可用凭证。使用确切的参数名 `access_key_id` 和 `secret_access_key`;其他拼写会被静默忽略,导致签名失败。 + +### 写入时出现连接或 DNS 错误 + +确认端点包含协议和端口,且应用可以访问它。Compose 网络内主机名为 `rustfs`;宿主机上使用 `http://localhost:9000`。 + +## 后续步骤 + +- 在采用更多 OpenDAL 操作之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [OpenDAL 文档](https://opendal.apache.org/docs/)在 Rust、Java 或 Node.js 中使用相同的 Operator。 diff --git a/content/zh/developer/integration/big-data/zeppelin.md b/content/zh/developer/integration/big-data/zeppelin.md new file mode 100644 index 00000000..3d28946c --- /dev/null +++ b/content/zh/developer/integration/big-data/zeppelin.md @@ -0,0 +1,109 @@ +--- +title: "Apache Zeppelin" +description: "通过 S3 notebook 仓库把 Apache Zeppelin 笔记本存储到 RustFS。" +--- + +本指南通过 Zeppelin 的 S3 笔记本存储,将数据分析笔记本 [Apache Zeppelin](https://github.com/apache/zeppelin) 连接到 **RustFS**。你将启动指向 RustFS 存储桶的 Zeppelin,创建一个笔记本,并验证笔记本文件已存储在 RustFS 中且在重启后依然保留。整个流程使用 `apache/zeppelin:0.12.0` 和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Browser["Browser"] -->|"notebook edits"| Z["Zeppelin :8080"] + Z -->|".zpln files"| RustFS["RustFS :9000"] +``` + +使用 `S3NotebookRepo` 存储类后,每个笔记本都会以 `.zpln` JSON 文件的形式持久化到存储桶的 `user/notebook/` 前缀下。Zeppelin 直接读写存储桶,因此笔记本可在容器重启后保留,还能在多个实例间共享。 + +## 1. 启动 Zeppelin + +创建环境文件,并替换两个凭证占位符: + +```ini title=".env" +AWS_ACCESS_KEY_ID= +AWS_SECRET_ACCESS_KEY= +``` + +在与 RustFS 相同的 Docker 网络中启动 Zeppelin,并带上 S3 存储设置: + +```bash +docker run -d --name zeppelin --network oo-rustfs_default \ + -p 8080:8080 \ + -e AWS_ACCESS_KEY_ID \ + -e AWS_SECRET_ACCESS_KEY \ + -e ZEPPELIN_NOTEBOOK_STORAGE=org.apache.zeppelin.notebook.repo.S3NotebookRepo \ + -e ZEPPELIN_NOTEBOOK_S3_BUCKET=my-bucket \ + -e ZEPPELIN_NOTEBOOK_S3_ENDPOINT=http://rustfs:9000 \ + -e ZEPPELIN_NOTEBOOK_S3_PATH_STYLE_ACCESS=true \ + apache/zeppelin:0.12.0 +``` + +Zeppelin 把 `ZEPPELIN_*` 环境变量作为配置属性读取,因此无需修改 `zeppelin-site.xml`。本指南使用现有的 `my-bucket`;笔记本会落到它的 `user/notebook/` 前缀下,该前缀由 S3 存储按需创建。 + +## 2. 创建笔记本 + +等待 `http://localhost:8080` 的 UI 就绪,在笔记本列表中创建名为 `rustfs-demo` 的笔记本并添加段落,或者使用 REST API: + +```bash +NOTE=$(curl -s -X POST "http://localhost:8080/api/notebook" \ + -H "Content-Type: application/json" \ + -d '{"name": "rustfs-demo"}' | python3 -c "import json,sys; print(json.load(sys.stdin)['body'])") +echo "note id: $NOTE" +``` + +## 3. 在 RustFS 中验证笔记本 + +列出笔记本前缀: + +```bash +rc ls rustfs/my-bucket/user/notebook/ -r +``` + +笔记本以"笔记本名 + ID"命名的 JSON 文件存储: + +```text +user/notebook/rustfs-demo_2N4PY7UY5.zpln +``` + +![RustFS 控制台中存储的 Zeppelin 笔记本](./images/rustfs-zeppelin-notebook.png) + +由于 Zeppelin 从存储桶加载笔记本,重启后依然存在: + +```bash +docker restart zeppelin +curl -s "http://localhost:8080/api/notebook" | head -c 200 +``` + +重启后笔记本列表中再次出现 `2N4PY7UY5`。 + +## 4. 停止或重置部署 + +停止 Zeppelin 并保留笔记本: + +```bash +docker rm -f zeppelin +``` + +笔记本保留在 `my-bucket` 的 `user/notebook/` 前缀下。若要删除它们,请移除该前缀: + +```bash +rc rm rustfs/my-bucket/user/notebook/ --recursive --force +``` + +## 故障排查 + +### Zeppelin 启动了但桶里始终没有笔记本 + +确认三个 `ZEPPELIN_NOTEBOOK_S3_*` 变量都已设置,且凭证环境变量能进入容器——S3 仓库在启动时初始化,任何修改后都需要重建容器。 + +### `UnknownHostException: my-bucket.rustfs` + +path-style 开关没有生效。保持 `ZEPPELIN_NOTEBOOK_S3_PATH_STYLE_ACCESS=true` 原样;禁用 path-style 后 Zeppelin 会把桶当作主机名。 + +## 后续步骤 + +- 在采用更多 Zeppelin 操作之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Zeppelin 存储文档](https://zeppelin.apache.org/docs/latest/setup/storage/storage.html#notebook-storage-in-s3)按用户组织笔记本前缀。 diff --git a/content/zh/developer/integration/devops/images/rustfs-jenkins-artifacts.png b/content/zh/developer/integration/devops/images/rustfs-jenkins-artifacts.png new file mode 100644 index 00000000..cd211cab Binary files /dev/null and b/content/zh/developer/integration/devops/images/rustfs-jenkins-artifacts.png differ diff --git a/content/zh/developer/integration/devops/index.md b/content/zh/developer/integration/devops/index.md index dd023482..a63e87bd 100644 --- a/content/zh/developer/integration/devops/index.md +++ b/content/zh/developer/integration/devops/index.md @@ -9,6 +9,7 @@ description: "通过 S3 兼容的对象存储接口,将 DevOps 平台与基础 - [Elasticsearch](./elasticsearch.md) - [Gitea](./gitea.md) +- [Jenkins](./jenkins.md) - [Terraform](./terraform.md) 请使用专用的存储桶保存制品、状态与遥测数据,并为凭证仅授予所需桶操作的权限。 diff --git a/content/zh/developer/integration/devops/jenkins.md b/content/zh/developer/integration/devops/jenkins.md new file mode 100644 index 00000000..8e604ebe --- /dev/null +++ b/content/zh/developer/integration/devops/jenkins.md @@ -0,0 +1,144 @@ +--- +title: "Jenkins" +description: "使用 Artifact Manager on S3 插件把 Jenkins 构建工件存储到 RustFS。" +--- + +本指南通过 Artifact Manager on S3 插件,将自动化服务器 [Jenkins](https://github.com/jenkinsci/jenkins) 连接到 **RustFS**。你将启动带该插件的 Jenkins,把工件管理器指向一个 RustFS 存储桶,运行一个归档工件的任务,并验证工件已存储在 RustFS 中。整个流程使用 `jenkins/jenkins:lts-jdk17`(Jenkins 2.5xx LTS)和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Dev["Developer"] -->|"trigger build"| J["Jenkins :8080"] + J -->|"archiveArtifacts"| RustFS["RustFS :9000"] +``` + +插件激活后,所有通过标准 `archiveArtifacts` 步骤(或 `stash`/`unstash`)发布工件的作业,都会把工件上传到 `jenkins-artifacts` 存储桶的配置前缀下,而不是存在控制器磁盘上。 + +## 1. 创建项目文件 + +先创建存储桶——插件会校验桶但不会创建它: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/jenkins-artifacts +``` + +创建一个在 Jenkins LTS 镜像中安装插件的 Dockerfile: + +```dockerfile title="Dockerfile" +FROM jenkins/jenkins:lts-jdk17 +USER root +RUN jenkins-plugin-cli --plugins artifact-manager-s3 aws-credentials +USER jenkins +``` + +`artifact-manager-s3` 插件会自动带上所需的 AWS 凭证支持;显式列出 `aws-credentials` 可以确保凭证类型可用。 + +构建镜像,并在与 RustFS 相同的 Docker 网络中启动 Jenkins: + +```bash +docker build -t jenkins-rustfs . +docker run -d --name jenkins --network oo-rustfs_default \ + -p 8080:8080 -v jenkins-home:/var/jenkins_home jenkins-rustfs +``` + +完成安装向导后创建 AWS 凭证:**Manage Jenkins → Credentials → global → Add Credentials**,类型选择 **AWS Credential**,ID 填 `rustfs-creds`,填入你的 RustFS 访问密钥和秘密密钥。 + +## 2. 配置工件管理器 + +打开 **Manage Jenkins → AWS Configuration**(来自 `aws-global-configuration` 插件)并设置: + +- **Region name**:`us-east-1` +- **Credentials**:`rustfs-creds` + +打开 **Manage Jenkins → System**,找到 **Artifact Management for Builds** 段。选择 **Cloud Provider Amazon S3**,并填写 S3 配置: + +- **S3 Bucket Name**:`jenkins-artifacts` +- **S3 Bucket Region**:`us-east-1` +- **Base Prefix**:`artifacts/` +- **Custom Endpoint**:`:9000`(Compose 网络内填 `rustfs:9000`) +- **Custom Signing Region**:`us-east-1` +- **Use Path Style URL**:启用 +- **Use Insecure HTTP**:启用 +- **Disable Session Token**:启用 + +非 AWS 且无 TLS 的端点必须使用 path-style 寻址和纯 HTTP。禁用会话令牌可阻止插件调用 AWS STS——静态访问密钥对无法响应 STS 调用。点击 **Validate S3 Bucket configuration** 确认配置无误后保存。 + +## 3. 运行归档工件的作业 + +创建一个生成文件并归档的 freestyle 作业(或流水线): + +```groovy +pipeline { + agent any + stages { + stage('Build') { + steps { + sh 'echo "jenkins artifact stored on rustfs" > report.txt' + } + } + } + post { + always { + archiveArtifacts 'report.txt' + } + } +} +``` + +运行构建并等待完成。工件上传对 RustFS 完全透明——作业配置里完全不出现 S3。 + +## 4. 在 RustFS 中验证对象 + +列出存储桶: + +```bash +rc ls rustfs/jenkins-artifacts/ -r +``` + +工件存储在前缀下,按作业和构建号组织: + +```text +artifacts/s3-artifacts-demo/3/artifacts/report.txt +``` + +![RustFS 控制台中存储的 Jenkins 工件](./images/rustfs-jenkins-artifacts.png) + +在构建页面下载工件时,读取的就是 RustFS 中的对象。 + +## 5. 停止或重置部署 + +停止 Jenkins 并保留数据: + +```bash +docker rm -f jenkins +``` + +工件保留在 `jenkins-artifacts` 存储桶中。若要删除它们,请移除存储桶: + +```bash +rc rb rustfs/jenkins-artifacts --force +``` + +## 故障排查 + +### `StsException: The security token included in the request is invalid` + +插件正在调用 AWS STS 获取会话凭证。在 S3 配置中启用 **Disable Session Token**——静态访问密钥对无法响应 STS 调用。 + +### `UnknownHostException: jenkins-artifacts.rustfs` + +插件使用了 virtual-hosted 寻址。启用 **Use Path Style URL**——RustFS 从 URL 路径而不是主机名解析桶。 + +### `No valid session credentials` 或空的凭证错误 + +确认 AWS Configuration 页面已选择凭证并**先于**使用 S3 存储桶设置保存,且凭证 ID 与你创建的一致。 + +## 后续步骤 + +- 在采用更多 Jenkins 集成之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Artifact Manager on S3 插件文档](https://plugins.jenkins.io/artifact-manager-s3/)了解 stash 支持与清理选项。 diff --git a/content/zh/developer/integration/devops/meta.json b/content/zh/developer/integration/devops/meta.json index 61ad9fb8..67c84c6e 100644 --- a/content/zh/developer/integration/devops/meta.json +++ b/content/zh/developer/integration/devops/meta.json @@ -3,6 +3,7 @@ "pages": [ "elasticsearch", "gitea", + "jenkins", "terraform" ] } diff --git a/content/zh/developer/integration/index.md b/content/zh/developer/integration/index.md index 82b9101f..f455c86d 100644 --- a/content/zh/developer/integration/index.md +++ b/content/zh/developer/integration/index.md @@ -9,10 +9,10 @@ description: "将 RustFS 与反向代理、备份工具、数据分析系统、 - [反向代理](./reverse-proxy/index.md)涵盖 Nginx、Traefik、Caddy 和 HAProxy。 - [备份](./backup/index.md)涵盖 Restic 和 Longhorn。 -- [数据分析](./big-data/index.md)涵盖 Iceberg。 -- [可观测性](./observability/index.md)涵盖 OpenObserve。 +- [数据分析](./big-data/index.md)涵盖 ClickHouse、Doris、Iceberg、Milvus、OpenDAL 和 Zeppelin 等数据分析系统。 +- [可观测性](./observability/index.md)涵盖 Fluentd、OpenObserve、OpenTelemetry、Thanos 和 Tempo 等遥测系统。 - [其他](./others/index.md)涵盖社区驱动的 Python capo SDK。 - [镜像仓库](./registry/index.md)涵盖 Harbor。 -- [DevOps](./devops/index.md)涵盖 Elasticsearch、Gitea 和 Terraform。 +- [DevOps](./devops/index.md)涵盖 Elasticsearch、Gitea、Jenkins 和 Terraform。 每篇指南都会说明配置集成系统时需要使用的 RustFS 端点和寻址要求。 \ No newline at end of file diff --git a/content/zh/developer/integration/observability/fluentd.md b/content/zh/developer/integration/observability/fluentd.md new file mode 100644 index 00000000..dad2ce93 --- /dev/null +++ b/content/zh/developer/integration/observability/fluentd.md @@ -0,0 +1,152 @@ +--- +title: "Fluentd" +description: "使用 S3 输出插件把 Fluentd 日志事件发送到 RustFS。" +--- + +本指南通过 `out_s3` 输出插件,将开源数据采集器 [Fluentd](https://github.com/fluent/fluentd) 连接到 **RustFS**。你将运行带 tail 输入源的 Fluentd,缓冲日志事件,并验证冲刷出的对象已存储在 RustFS 中。整个流程使用 `fluent/fluentd:v1.17-1`、`fluent-plugin-s3` 1.8.6 和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + App["Application"] -->|"writes lines"| File["app.log"] + File -->|tail| Fluentd["Fluentd"] + Fluentd -->|"gzip objects"| RustFS["RustFS :9000"] +``` + +tail 输入源读取 `app.log` 的新增行并交给 S3 输出,输出按时间键缓冲事件,每个冲刷窗口上传一个 gzip 对象。 + +## 1. 创建项目文件 + +先创建存储桶——Fluentd 不会创建桶: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/fluentd-data +``` + +官方 Fluentd 镜像以非 root 用户运行,无法在启动时安装 gem,因此构建一个带 S3 插件的小镜像: + +```dockerfile title="Dockerfile" +FROM fluent/fluentd:v1.17-1 +USER root +RUN gem install fluent-plugin-s3 --no-document +USER fluent +``` + +创建 Fluentd 配置,并替换两个凭证占位符: + +```nginx title="fluent.conf" + + @type tail + path /var/log/app.log + pos_file /var/log/app.log.pos + tag rustfs.demo + + @type none + + + + + @type s3 + aws_key_id + aws_sec_key + s3_bucket fluentd-data + s3_endpoint http://rustfs:9000/ + s3_region us-east-1 + force_path_style true + path fluentd-logs + + @type memory + timekey 30s + timekey_wait 0s + flush_mode immediate + + +``` + +`force_path_style true` 是必需的——缺少它插件会构造 `fluentd-data.rustfs` 作为主机名,所有请求都会因 DNS 错误失败。Compose 网络内主机名为 `rustfs`;宿主机上使用 (`http://localhost:9000/`)。 + +构建镜像,并在与 RustFS 相同的 Docker 网络中启动 Fluentd: + +```bash +docker build -t fluentd-rustfs . +mkdir -p logs +docker run -d --name fluentd --network oo-rustfs_default \ + -v "$PWD/fluent.conf":/fluentd/etc/fluent.conf:ro \ + -v "$PWD/logs":/var/log fluentd-rustfs +``` + +## 2. 产生日志事件 + +向被监视的文件追加内容: + +```bash +echo "rustfs fluentd demo line 1" >> logs/app.log +echo "rustfs fluentd demo line 2" >> logs/app.log +``` + +在 30 秒时间键和立即冲刷模式下,每个窗口关闭后不久就会上传一个 gzip 对象。等待约一分钟。 + +## 3. 在 RustFS 中验证对象 + +列出存储桶: + +```bash +rc ls rustfs/fluentd-data/ -r +``` + +每个冲刷窗口产生一个 gzip 对象: + +```text +fluentd-logs20260921154400_0.gz +fluentd-logs20260921154400_1.gz +``` + +读回一个对象确认事件完整: + +```bash +rc cat rustfs/fluentd-data/fluentd-logs20260921154400_0.gz | gunzip +``` + +![RustFS 控制台中存储的 Fluentd 日志对象](./images/rustfs-fluentd-logs.png) + +## 4. 停止或重置部署 + +停止 Fluentd 并保留数据: + +```bash +docker rm -f fluentd +``` + +对象保留在 `fluentd-data` 存储桶中。若要删除它们,请移除存储桶: + +```bash +rc rb rustfs/fluentd-data --force +``` + +## 故障排查 + +### `Unknown output plugin 's3'` + +插件未安装。确认 Dockerfile 在以 `root` 身份安装 `fluent-plugin-s3` 后才切回 `fluent` 用户——以默认用户在容器启动时安装会因权限错误失败。 + +### `Failed to open TCP connection to fluentd-data.rustfs` + +正在使用 virtual-hosted 寻址。在 `s3` 输出中添加 `force_path_style true`,让桶保留在 URL 路径中。 + +### worker 反复崩溃重启 + +桶不存在时输出会硬失败。启动 Fluentd 前先创建 `fluentd-data`,并检查启动日志: + +```bash +docker logs fluentd | grep -iE "error|bucket" | tail +``` + +## 后续步骤 + +- 在采用更多 Fluentd 输出之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [fluent-plugin-s3 文档](https://github.com/fluent/fluent-plugin-s3)了解对象键格式与压缩选项。 diff --git a/content/zh/developer/integration/observability/images/rustfs-fluentd-logs.png b/content/zh/developer/integration/observability/images/rustfs-fluentd-logs.png new file mode 100644 index 00000000..eeebccad Binary files /dev/null and b/content/zh/developer/integration/observability/images/rustfs-fluentd-logs.png differ diff --git a/content/zh/developer/integration/observability/images/rustfs-otel-logs.png b/content/zh/developer/integration/observability/images/rustfs-otel-logs.png new file mode 100644 index 00000000..b9fb0c30 Binary files /dev/null and b/content/zh/developer/integration/observability/images/rustfs-otel-logs.png differ diff --git a/content/zh/developer/integration/observability/index.md b/content/zh/developer/integration/observability/index.md index 5d85e502..8b6a450b 100644 --- a/content/zh/developer/integration/observability/index.md +++ b/content/zh/developer/integration/observability/index.md @@ -7,7 +7,9 @@ description: "通过 S3 兼容对象存储接口,将可观测性平台连接 ## 平台 +- [Fluentd](./fluentd.md) - [OpenObserve](./openobserve.md) +- [OpenTelemetry](./opentelemetry.md) - [Loki](./loki.md) - [Tempo](./tempo.md) - [Thanos](./thanos.md) diff --git a/content/zh/developer/integration/observability/meta.json b/content/zh/developer/integration/observability/meta.json index 4db188bf..0acec73a 100644 --- a/content/zh/developer/integration/observability/meta.json +++ b/content/zh/developer/integration/observability/meta.json @@ -1,8 +1,10 @@ { "title": "可观测性", "pages": [ - "openobserve", + "fluentd", "loki", + "openobserve", + "opentelemetry", "tempo", "thanos" ] diff --git a/content/zh/developer/integration/observability/opentelemetry.md b/content/zh/developer/integration/observability/opentelemetry.md new file mode 100644 index 00000000..4a8ff46c --- /dev/null +++ b/content/zh/developer/integration/observability/opentelemetry.md @@ -0,0 +1,145 @@ +--- +title: "OpenTelemetry" +description: "通过 AWS S3 导出器把 OpenTelemetry Collector 日志导出到 RustFS。" +--- + +本指南通过 `awss3` 导出器,将 [OpenTelemetry Collector](https://opentelemetry.io/docs/collector/)——CNCF 遥测管道——连接到 **RustFS**。你将运行带 `filelog` 接收器的 contrib 版 Collector,把一个日志文件的内容发送到 `otel-data` 存储桶,并验证 RustFS 中的分区对象。整个流程使用 `otel/opentelemetry-collector-contrib:0.138.0` 和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Log["Application log file"] -->|filelog| Collector["OTel Collector"] + Collector -->|awss3| RustFS["RustFS :9000"] +``` + +`filelog` 接收器监视日志文件,`awss3` 导出器把批次上传到存储桶并按时间分区。指标和链路可以通过同一导出器走各自的管道。 + +## 1. 创建项目文件 + +先创建存储桶——导出器不会创建桶: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/otel-data +``` + +创建 Collector 配置,并替换两个凭证占位符: + +```yaml title="config.yaml" +receivers: + filelog: + include: [/var/log/app.log] + start_at: beginning + +exporters: + awss3: + s3uploader: + region: us-east-1 + s3_bucket: otel-data + endpoint: http://rustfs:9000 + s3_force_path_style: true + disable_ssl: true + file_prefix: logs/app + marshaler: body + +service: + pipelines: + logs: + receivers: [filelog] + exporters: [awss3] +``` + +导出器从标准的 `AWS_ACCESS_KEY_ID`、`AWS_SECRET_ACCESS_KEY`、`AWS_REGION` 环境变量读取凭证。`marshaler: body` 把每行日志写为纯文本;省略它则存储 OTLP JSON。纯 HTTP 的非 AWS 端点必须设置 `s3_force_path_style` 和 `disable_ssl`。 + +创建 Compose 文件: + +```yaml title="compose.yaml" +services: + collector: + image: otel/opentelemetry-collector-contrib:0.138.0 + command: ["--config=/etc/otelcol-contrib/config.yaml"] + environment: + AWS_ACCESS_KEY_ID: + AWS_SECRET_ACCESS_KEY: + AWS_REGION: us-east-1 + volumes: + - ./config.yaml:/etc/otelcol-contrib/config.yaml:ro + - ./app.log:/var/log/app.log + networks: + - otel + +networks: + otel: +``` + +## 2. 启动 Collector 并产生日志 + +启动整个栈并向被监视的文件追加内容: + +```bash +echo "otel demo log line one" > app.log +docker compose up -d +echo "second line after start" >> app.log +``` + +Collector 从头读取文件,在分区滚动时上传每个缓冲区,因此最后一条日志之后请等待约一分钟再检查。 + +## 3. 在 RustFS 中验证对象 + +列出存储桶: + +```bash +rc ls rustfs/otel-data/ -r +``` + +日志记录按时间分区加上配置的文件前缀上传: + +```text +year=2026/month=09/day=21/hour=15/minute=53/logs/applogs_288361608.txt +``` + +读回一个对象确认日志行完整到达: + +```bash +rc cat rustfs/otel-data/year=2026/month=09/day=21/hour=15/minute=53/logs/applogs_288361608.txt +``` + +```text +otel demo log line one +second line after start +``` + +![RustFS 控制台中存储的 OpenTelemetry 日志对象](./images/rustfs-otel-logs.png) + +## 4. 停止或重置部署 + +停止 Collector 并保留数据: + +```bash +docker compose down +``` + +对象保留在 `otel-data` 存储桶中。若要删除它们,请移除存储桶: + +```bash +rc rb rustfs/otel-data --force +``` + +## 故障排查 + +### Collector 启动时报 `has invalid keys` + +`awss3` 导出器的配置结构随 Collector 版本变化。0.138 版把上传配置嵌套在 `s3uploader` 下(如上所示),端点键名为 `endpoint`(不是 `s3_endpoint`);更新的版本把这些键移到顶层。请核对你所用版本的 README。 + +### 存储桶中没有出现对象 + +确认 Collector 容器设置了凭证环境变量、`s3_force_path_style` 为 `true`,并且 `http://rustfs:9000` 在 Compose 网络内可达。在同一管道上启用 `debug` 导出器可以确认数据是否在流动。 + +## 后续步骤 + +- 在采用更多 Collector 操作之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [AWS S3 导出器文档](https://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/exporter/awss3exporter)添加指标和链路管道。