diff --git a/content/de/developer/integration/ai/images/rustfs-ray-data.png b/content/de/developer/integration/ai/images/rustfs-ray-data.png new file mode 100644 index 00000000..394f4499 Binary files /dev/null and b/content/de/developer/integration/ai/images/rustfs-ray-data.png differ diff --git a/content/de/developer/integration/ai/index.md b/content/de/developer/integration/ai/index.md new file mode 100644 index 00000000..bad31e8a --- /dev/null +++ b/content/de/developer/integration/ai/index.md @@ -0,0 +1,12 @@ +--- +title: "AI" +description: "Verbinden Sie KI-Plattformen über S3-kompatible Objektspeicher-Schnittstellen mit RustFS." +--- + +Nutzen Sie **RustFS** als Objektspeicher-Layer für KI- und Machine-Learning-Plattformen, die einen S3-kompatiblen Endpunkt unterstützen. + +## Plattformen + +- [Ray](./ray.md) + +Speichern Sie Trainingsdaten und Checkpoints in dedizierten Buckets und beschränken Sie die Anmeldeinformationen auf die erforderlichen Bucket-Operationen. diff --git a/content/de/developer/integration/ai/meta.json b/content/de/developer/integration/ai/meta.json new file mode 100644 index 00000000..573fc520 --- /dev/null +++ b/content/de/developer/integration/ai/meta.json @@ -0,0 +1,6 @@ +{ + "title": "AI", + "pages": [ + "ray" + ] +} diff --git a/content/de/developer/integration/ai/ray.md b/content/de/developer/integration/ai/ray.md new file mode 100644 index 00000000..e9f71b94 --- /dev/null +++ b/content/de/developer/integration/ai/ray.md @@ -0,0 +1,103 @@ +--- +title: "Ray" +description: "Use Ray Data with RustFS as S3-compatible storage for dataset writes and reads." +--- + +This guide connects [Ray](https://github.com/ray-project/ray) — the distributed AI and Python compute framework — to **RustFS** through Ray Data's S3 filesystem support. You will run a Ray job inside the official image, write a dataset as Parquet to a RustFS bucket, read it back, and verify the objects. The workflow was verified with `rayproject/ray:2.44.0-py311` (Ray 2.44, pyarrow filesystem) and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Job["Ray job"] -->|"ray.data"| DS["Dataset"] + DS -->|"Parquet files"| RustFS["RustFS :9000"] +``` + +Ray Data reads and writes datasets through pyarrow's `S3FileSystem`. Passing an `S3FileSystem` configured for RustFS redirects every dataset operation — Parquet, CSV, JSON — to the bucket. + +## 1. Create the job file + +Create the script, replacing all connection placeholders: + +```python title="ray_s3.py" +import ray +ray.init(ignore_reinit_error=True) + +import pandas as pd +from pyarrow.fs import S3FileSystem + +fs = S3FileSystem( + endpoint_override="http://:9000", + access_key="", + secret_key="", + region="us-east-1", +) + +df = pd.DataFrame({"id": range(5), "value": [x * 1.5 for x in range(5)]}) +ds = ray.data.from_pandas(df) +ds.write_parquet("my-bucket/ray-demo/events/", filesystem=fs) + +back = ray.data.read_parquet("my-bucket/ray-demo/events/", filesystem=fs).take_all() +print("rows:", len(back)) +print("sample:", back[0]) +ray.shutdown() +``` + +`endpoint_override` takes the full endpoint URL including the scheme. pyarrow's `S3FileSystem` uses path-style requests for custom endpoints, so no extra flag is needed. The same filesystem object works for `write_csv`, `read_json`, and the other Ray Data methods. + +## 2. Run the job + +Run the script in the Ray image on the same Docker network as RustFS: + +```bash +docker run --rm --network oo-rustfs_default \ + -v "$PWD/ray_s3.py":/tmp/ray_s3.py \ + rayproject/ray:2.44.0-py311 python /tmp/ray_s3.py +``` + +```text +rows: 5 +sample: {'id': 0, 'value': 0.0} +``` + +## 3. Verify objects in RustFS + +List the dataset prefix: + +```bash +rc ls rustfs/my-bucket/ray-demo/ -r +``` + +Ray Data wrote the dataset as a Parquet block: + +```text +ray-demo/events/0_000000_000000.parquet +``` + +![Ray dataset files stored in the RustFS Console](./images/rustfs-ray-data.png) + +## 4. Stop or reset + +Ray Data holds no state of its own. To delete the demo dataset: + +```bash +rc rm rustfs/my-bucket/ray-demo/ --recursive --force +``` + +## Troubleshooting + +### `Unable to connect to endpoint` or timeouts + +Confirm `endpoint_override` includes the scheme and is reachable from the Ray container. Inside a Compose network the hostname is `rustfs`; from the host use `http://localhost:9000`. + +### `Access Denied` on write + +Confirm the access key and secret key are passed to `S3FileSystem` itself — Ray does not read the container's AWS environment variables through this code path. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Ray operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Ray Data documentation](https://docs.ray.io/en/latest/data/data.html) to chain transformations, training ingestion, and checkpointing on the same bucket. diff --git a/content/de/developer/integration/backup/images/rustfs-kopia-repo.png b/content/de/developer/integration/backup/images/rustfs-kopia-repo.png new file mode 100644 index 00000000..4241891c Binary files /dev/null and b/content/de/developer/integration/backup/images/rustfs-kopia-repo.png differ diff --git a/content/de/developer/integration/backup/images/rustfs-velero-backups.png b/content/de/developer/integration/backup/images/rustfs-velero-backups.png new file mode 100644 index 00000000..c3e9b169 Binary files /dev/null and b/content/de/developer/integration/backup/images/rustfs-velero-backups.png differ diff --git a/content/de/developer/integration/backup/index.md b/content/de/developer/integration/backup/index.md index de236c2a..8e54d4cf 100644 --- a/content/de/developer/integration/backup/index.md +++ b/content/de/developer/integration/backup/index.md @@ -8,6 +8,8 @@ Verwenden Sie **RustFS** als Objektspeicher-Backend für Backup-Tools, die Repos ## Systeme - [Restic](./restic.md) +- [Velero](./velero.md) +- [Kopia](./kopia.md) - [Longhorn](./longhorn.md) Halten Sie Backup-Jobs in einem eigenen Bucket und Präfix und verwenden Sie Anmeldedaten, die auf die erforderlichen Bucket-Operationen beschränkt sind. \ No newline at end of file diff --git a/content/de/developer/integration/backup/kopia.md b/content/de/developer/integration/backup/kopia.md new file mode 100644 index 00000000..fed21e5a --- /dev/null +++ b/content/de/developer/integration/backup/kopia.md @@ -0,0 +1,125 @@ +--- +title: "Kopia" +description: "Back up files to RustFS with Kopia's S3 repository backend." +--- + +This guide connects [Kopia](https://github.com/kopia/kopia) — the open-source backup and restore tool — to **RustFS** as an S3 repository. You will create a repository in a RustFS bucket, take a snapshot of a directory, restore it into an empty directory, and compare checksums. The workflow was verified with `kopia/kopia:0.18.1` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Source["Source files"] -->|snapshot| Kopia["Kopia"] + Kopia -->|"encrypted blocks"| RustFS["RustFS :9000"] + Kopia -->|restore| Restore["Restored files"] +``` + +Kopia stores the repository format files and deduplicated, encrypted content blocks in the bucket. Restores read the blocks back and reassemble the original files, so the checksum of every restored file must match the source. + +## 1. Create the repository + +Create the bucket first, then initialize the Kopia repository inside it. The endpoint is a bare `host:port` (no scheme); `--disable-tls` switches the client to plain HTTP: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/kopia-backups + +docker run --rm --network oo-rustfs_default kopia/kopia:0.18.1 repository create s3 \ + --bucket kopia-backups \ + --access-key \ + --secret-access-key \ + --endpoint :9000 \ + --region us-east-1 \ + --disable-tls \ + --password \ + --override-username demo --override-hostname workstation +``` + +Kopia validates the provider by reading and writing through the S3 API before it reports success. + +## 2. Connect, snapshot, and restore + +Run the following commands from the directory holding your `repository.config` (created by the previous step). Kopia reads the connection settings from that file, so the S3 flags are only needed once: + +```bash +export KOPIA_PASSWORD= +export KOPIA_CONFIG_PATH=/config/repository.config + +alias kopia='docker run --rm --network oo-rustfs_default \ + -e KOPIA_PASSWORD -e KOPIA_CONFIG_PATH \ + -v "$PWD/config:/config" \ + -v "$PWD/source:/source:ro" \ + -v "$PWD/restore:/restore" kopia/kopia:0.18.1' + +kopia repository connect s3 \ + --bucket kopia-backups \ + --access-key \ + --secret-access-key \ + --endpoint :9000 \ + --region us-east-1 --disable-tls \ + --override-username demo --override-hostname workstation + +kopia snapshot create /source +kopia snapshot list +``` + +Restore the snapshot into an empty directory and compare checksums with the source. The snapshot ID is the `ka...` identifier printed by `snapshot list`: + +```bash +kopia restore /restore + +sha256sum source/blob.bin restore/blob.bin +``` + +```text +bec4530e2798465b... source/blob.bin +bec4530e2798465b... restore/blob.bin +``` + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/kopia-backups/ -r +``` + +The output shows the repository format files plus the packed content blocks written by the snapshot: + +```text +kopia.blobcfg +kopia.repository +p0000.../... +``` + +![Kopia repository blocks stored in the RustFS Console](./images/rustfs-kopia-repo.png) + +## 4. Stop or reset + +Kopia is a client-side tool and holds no running state. To delete the repository and all snapshots, remove the bucket: + +```bash +rc rb rustfs/kopia-backups --force +``` + +## Troubleshooting + +### `Endpoint url cannot have fully qualified paths` + +The endpoint must be a bare `host:port` value without a scheme or path — Kopia builds the object URLs itself. + +### `server gave HTTP response to HTTPS client` + +Without `--disable-tls`, Kopia speaks HTTPS. RustFS without TLS needs the `--disable-tls` flag on both `repository create s3` and `repository connect s3`. + +### `can't connect to storage` with a DNS error + +The S3 client is using virtual-hosted addressing. Keep the endpoint as a bare `host:port` value; Kopia uses path-style requests for endpoints given in that form. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Kopia operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Kopia repository documentation](https://kopia.io/docs/repositories/) to add policies, retention, and scheduled snapshots. diff --git a/content/de/developer/integration/backup/meta.json b/content/de/developer/integration/backup/meta.json index 3ea20cbe..dac5156c 100644 --- a/content/de/developer/integration/backup/meta.json +++ b/content/de/developer/integration/backup/meta.json @@ -1,7 +1,9 @@ { "title": "Backup", "pages": [ + "kopia", + "longhorn", "restic", - "longhorn" + "velero" ] } diff --git a/content/de/developer/integration/backup/velero.md b/content/de/developer/integration/backup/velero.md new file mode 100644 index 00000000..0140296b --- /dev/null +++ b/content/de/developer/integration/backup/velero.md @@ -0,0 +1,153 @@ +--- +title: "Velero" +description: "Back up Kubernetes cluster resources to RustFS with Velero's AWS object store provider." +--- + +This guide connects [Velero](https://github.com/vmware-tanzu/velero) — the Kubernetes backup and restore tool — to **RustFS** as its object storage backend. You will install Velero into a Kubernetes cluster with a RustFS backup location, back up cluster resources, delete them, restore from the bucket, and verify the restored objects. The workflow was verified with Velero CLI 1.16.2, `velero-plugin-for-aws:v1.12.2` on k3s (Kubernetes 1.30), and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need a Kubernetes cluster with `kubectl` access and the Velero CLI. This guide is intended for integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + K8s["Kubernetes cluster"] -->|"resources"| V["Velero"] + V -->|"backups + logs"| RustFS["RustFS :9000"] +``` + +Velero serializes Kubernetes resources and (optionally) pod volume data into gzip archives under `backups//` in the bucket. Restores download those archives and recreate the resources in the cluster. + +## 1. Create the backup bucket + +Create a dedicated bucket — Velero does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/velero-backups +``` + +## 2. Install Velero + +Write the credentials to a file and install the Velero server. The `s3Url` must be an endpoint reachable from the cluster pods themselves — use the node IP or an internal address, not a port forward from your workstation: + +```bash +cat > velero-creds < +aws_secret_access_key= +EOF + +kubectl create namespace velero +kubectl create secret generic cloud-credentials \ + --namespace velero --from-file=cloud=velero-creds + +velero install \ + --provider aws \ + --plugins velero/velero-plugin-for-aws:v1.12.2 \ + --bucket velero-backups \ + --backup-location-config region=us-east-1,s3ForcePathStyle="true",s3Url=http://:9000 \ + --secret-file velero-creds \ + --use-volume-snapshots=false +``` + +Wait for the deployment to become ready and the backup location to turn `Available`: + +```bash +kubectl -n velero get pods +velero backup-location get +``` + +```text +NAME PROVIDER BUCKET/PREFIX PHASE LAST VALIDATED ACCESS MODE DEFAULT +default aws velero-backups Available 2026-09-22 ... ReadWrite true +``` + +## 3. Back up cluster resources + +Create two demo resources and back up the whole default namespace scope: + +```bash +kubectl create configmap demo-cm --from-literal=key=rustfs-velero-demo +kubectl create deployment nginx --image=nginx:1.27 + +velero backup create demo-backup --wait +velero backup get +``` + +The backup uploads the resource archives and logs to RustFS. + +## 4. Verify the backup in RustFS + +List the bucket: + +```bash +rc ls rustfs/velero-backups/ -r +``` + +The backup archive set is stored under the `backups/` prefix: + +```text +backups/demo-backup/demo-backup-resources.json.gz +backups/demo-backup/demo-backup-logs.gz +backups/demo-backup/demo-backup-itemoperations.json.gz +``` + +![Velero backups stored in the RustFS Console](./images/rustfs-velero-backups.png) + +## 5. Restore and verify + +Delete the demo resources and restore them from the backup. The restore reads the archives from RustFS and recreates the resources: + +```bash +kubectl delete configmap demo-cm +kubectl delete deployment nginx + +velero restore create --from-backup demo-backup --wait +``` + +Confirm the resources are back with their original content: + +```bash +kubectl get configmap demo-cm -o jsonpath="{.data.key}" +kubectl get deployment nginx +``` + +```text +rustfs-velero-demo +NAME READY UP-TO-DATE AVAILABLE AGE +nginx 1/1 1 1 10s +``` + +## 6. Stop or reset + +The backups stay in the `velero-backups` bucket and restore into any cluster that runs Velero against the same bucket. To remove the demo backup from RustFS: + +```bash +velero backup delete demo-backup --confirm +rc rm rustfs/velero-backups/backups/ --recursive --force +``` + +## Troubleshooting + +### The backup location stays `Unavailable` + +The Velero pod cannot reach the `s3Url`. Endpoints on a Docker custom bridge network are not routable from Kubernetes pods — use the Docker host bridge IP (for example `http://172.17.0.1:9000` on a single-host cluster) or a node address. Patch the location and let Velero revalidate: + +```bash +kubectl -n velero patch backupstoragelocation default --type merge -p \ + '{"spec":{"config":{"s3Url":"http://172.17.0.1:9000"}}}' +``` + +### The backup ends `PartiallyFailed` with a daemonset error + +The error `daemonset pod not found in running state` means the node-agent pod (used for pod volume backups) was not running. It does not affect the cluster-resource archives. Wait for the agent or keep `--default-volumes-to-fs-backup` only when the agent is running. + +### `FailedValidation` right after install + +The first backup attempt runs against a location that has not validated yet. Wait for `velero backup-location get` to report `Available` before creating backups. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Velero operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Velero documentation](https://velero.io/docs/main/) to add schedules, volume snapshots, and cluster migration flows. diff --git a/content/de/developer/integration/big-data/hudi.md b/content/de/developer/integration/big-data/hudi.md new file mode 100644 index 00000000..42026454 --- /dev/null +++ b/content/de/developer/integration/big-data/hudi.md @@ -0,0 +1,109 @@ +--- +title: "Apache Hudi" +description: "Write Apache Hudi tables to RustFS through Spark and the s3a connector." +--- + +This guide connects [Apache Hudi](https://github.com/apache/hudi) — the transactional data lake platform — to **RustFS** as its copy-on-write table store. You will run Spark with the Hudi bundle, write a table to the `s3a://` location inside a RustFS bucket, read it back, and verify the table files. The workflow was verified with `apache/spark:3.5.6`, `hudi-spark3.5-bundle_2.12:0.15.0`, `hadoop-aws:3.3.4`, and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Job["Spark + Hudi"] -->|"commit + parquet"| RustFS["RustFS :9000"] +``` + +Hudi stores the `.hoodie/` timeline, commit files, and Parquet data blocks under the table path in the bucket. Reads resolve the latest table snapshot from the timeline, so every write and query goes through the S3 API. + +## 1. Run the Spark write + +Hudi requires the Kryo serializer and the s3a credentials as Hadoop properties. The following `spark-shell` session creates a non-partitioned table in `my-bucket` — replace all connection placeholders: + +```scala +import org.apache.spark.sql.SaveMode + +val df = Seq((1, "a"), (2, "b"), (3, "c")).toDF("id", "name") +df.write.format("hudi") + .option("hoodie.table.name", "events") + .option("hoodie.datasource.write.recordkey.field", "id") + .option("hoodie.datasource.write.precombine.field", "name") + .option("hoodie.datasource.write.partitionpath.field", "") + .mode(SaveMode.Overwrite) + .save("s3a://my-bucket/hudi-demo/events") + +val back = spark.read.format("hudi").load("s3a://my-bucket/hudi-demo/events") +back.select("id", "name").show() +``` + +```bash +docker run --rm --network oo-rustfs_default \ + -v "$PWD/hudi_test.scala":/tmp/hudi_test.scala \ + apache/spark:3.5.6 /opt/spark/bin/spark-shell \ + --packages org.apache.hudi:hudi-spark3.5-bundle_2.12:0.15.0,org.apache.hadoop:hadoop-aws:3.3.4 \ + --conf spark.serializer=org.apache.spark.serializer.KryoSerializer \ + --conf spark.hadoop.fs.s3a.endpoint=http://rustfs:9000 \ + --conf spark.hadoop.fs.s3a.path.style.access=true \ + --conf spark.hadoop.fs.s3a.access.key= \ + --conf spark.hadoop.fs.s3a.secret.key= \ + --conf spark.hadoop.fs.s3a.region=us-east-1 \ + --conf spark.jars.ivy=/tmp/.ivy2 \ + --conf spark.sql.shuffle.partitions=2 \ + -i /tmp/hudi_test.scala +``` + +The `--packages` flags download the Hudi bundle and the s3a connector on first run. The `spark.jars.ivy` setting avoids Ivy cache permission errors in the container. The read-back `show()` prints the three rows: + +```text ++---+----+ +| 2| b| +| 3| c| +| 1| a| ++---+----+ +``` + +## 2. Verify objects in RustFS + +List the table prefix: + +```bash +rc ls rustfs/my-bucket/hudi-demo/ -r +``` + +The Hudi table layout appears under the table path — the `.hoodie/` timeline with the commit file, plus the Parquet data block: + +```text +hudi-demo/events/.hoodie/20260922015803291.commit +hudi-demo/events/.hoodie/20260922015803291.commit.requested +hudi-demo/events/// +``` + +![Hudi table files stored in the RustFS Console](./images/rustfs-hudi-table.png) + +## 3. Stop or reset + +Spark runs as a one-shot client and holds no state. To delete the demo table: + +```bash +rc rm rustfs/my-bucket/hudi-demo/ --recursive --force +``` + +## Troubleshooting + +### `hoodie only support org.apache.spark.serializer.KryoSerializer as spark.serializer` + +Hudi rejects the default Spark serializer. Pass `--conf spark.serializer=org.apache.spark.serializer.KryoSerializer` as shown above. + +### `Partition-path field has to be non-empty` or keygenerator class errors + +For a non-partitioned table, keep the default key generator (do not set `hoodie.datasource.write.keygenerator.class`) and set `hoodie.datasource.write.partitionpath.field` to the empty string. + +### `NoSuchMethodError` or `ClassNotFoundException` in the Hudi write + +The Spark and Hudi bundle versions must match: `hudi-spark3.5-bundle_2.12` goes with Spark 3.5.x, and the bundled Hadoop client must be compatible with `hadoop-aws:3.3.4`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Hudi operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Hudi Spark guide](https://hudi.apache.org/docs/quick-start-guide/) to add upserts, compaction, and query integrations on the same table. diff --git a/content/de/developer/integration/big-data/images/rustfs-hudi-table.png b/content/de/developer/integration/big-data/images/rustfs-hudi-table.png new file mode 100644 index 00000000..b91c9157 Binary files /dev/null and b/content/de/developer/integration/big-data/images/rustfs-hudi-table.png differ diff --git a/content/de/developer/integration/big-data/images/rustfs-lakefs-repo.png b/content/de/developer/integration/big-data/images/rustfs-lakefs-repo.png new file mode 100644 index 00000000..0236a479 Binary files /dev/null and b/content/de/developer/integration/big-data/images/rustfs-lakefs-repo.png differ diff --git a/content/de/developer/integration/big-data/index.md b/content/de/developer/integration/big-data/index.md index 342e0be6..b5850b9f 100644 --- a/content/de/developer/integration/big-data/index.md +++ b/content/de/developer/integration/big-data/index.md @@ -8,6 +8,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo ## Systems - [ClickHouse](./clickhouse.md) +- [Hudi](./hudi.md) - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) - [Milvus](./milvus.md) @@ -15,6 +16,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [OpenDAL](./opendal.md) - [DuckDB](./duckdb.md) - [Doris](./doris.md) +- [lakeFS](./lakefs.md) - [InfluxDB](./influxdb.md) - [Spark](./spark.md) - [Flink](./flink.md) diff --git a/content/de/developer/integration/big-data/lakefs.md b/content/de/developer/integration/big-data/lakefs.md new file mode 100644 index 00000000..ba63ef22 --- /dev/null +++ b/content/de/developer/integration/big-data/lakefs.md @@ -0,0 +1,160 @@ +--- +title: "lakeFS" +description: "Run lakeFS with RustFS as its S3 blockstore for versioned data lakes." +--- + +This guide connects [lakeFS](https://github.com/treeverse/lakeFS) — the Git-like data lake versioning layer — to **RustFS** as its S3 blockstore. You will start lakeFS, create a repository whose storage namespace points at a RustFS bucket, commit an object, and verify that the lakeFS metadata and data files live in RustFS. The workflow was verified with `treeverse/lakefs:1.58.0` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["lakectl / API"] -->|HTTP| LakeFS["lakeFS :8000"] + LakeFS -->|"metadata + data"| RustFS["RustFS :9000"] +``` + +lakeFS stores repository metadata and committed data files in the blockstore under the storage namespace. The bucket holds one `repo/` prefix containing `_lakefs/` metadata and `data/` objects; lakeFS reads and writes them through the S3 API. + +## 1. Create the project files + +Create the bucket, then the lakeFS configuration: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/lakefs-data +``` + +```yaml title="config.yaml" +database: + type: local + local: + path: /lakefs/data + +blockstore: + type: s3 + s3: + endpoint: http://rustfs:9000 + region: us-east-1 + force_path_style: true + +auth: + encrypt: + secret_key: change-me-to-a-random-string + +gateways: + s3: + domain_name: rustfs-gateway.local + +logging: + format: text + level: info +``` + +`force_path_style: true` is required — without it lakeFS builds `lakefs-data.` as a hostname and every request fails with a DNS error. The credentials are supplied through the standard `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` environment variables. + +Create the environment file for the service credentials and the initial admin user, replacing the placeholders: + +```ini title=".env" +AWS_ACCESS_KEY_ID= +AWS_SECRET_ACCESS_KEY= +LAKEFS_INSTALLATION_USER_NAME=admin +LAKEFS_INSTALLATION_ACCESS_KEY_ID= +LAKEFS_INSTALLATION_SECRET_ACCESS_KEY= +LAKEFS_STATS_ENABLED=false +``` + +## 2. Start lakeFS + +```bash +docker run -d --name lakefs --network oo-rustfs_default \ + -p 8000:8000 \ + -v "$PWD/config.yaml":/etc/lakefs/config.yaml:ro \ + --env-file .env \ + treeverse/lakefs:1.58.0 run +``` + +Wait for `http://localhost:8000/api/healthcheck` to return `200`, then create a repository with a storage namespace inside the bucket: + +```bash +curl -s -u : \ + -X POST http://localhost:8000/api/v1/repositories \ + -H "Content-Type: application/json" \ + -d '{"name": "rustfs-demo", "storage_namespace": "s3://lakefs-data/repo"}' +``` + +## 3. Commit an object + +Upload a file to the `main` branch and commit it: + +```bash +echo "hello from lakefs on rustfs" > hello.txt + +curl -s -u : \ + -X POST "http://localhost:8000/api/v1/repositories/rustfs-demo/branches/main/objects?path=hello.txt" \ + --data-binary @hello.txt + +curl -s -u : \ + -X POST "http://localhost:8000/api/v1/repositories/rustfs-demo/branches/main/commits" \ + -H "Content-Type: application/json" \ + -d '{"message": "add hello"}' +``` + +Read the object back through the branch ref — the response body is the committed content: + +```bash +curl -s -u : \ + "http://localhost:8000/api/v1/repositories/rustfs-demo/refs/main/objects?path=hello.txt" +``` + +## 4. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/lakefs-data/ -r +``` + +The output shows the lakeFS metadata objects and the committed data file under the repository prefix: + +```text +repo/_lakefs/19b2b26e37cb20fc6763c527f88eb5151891b04a2c8c9ddd32870c5c3f353281 +repo/data/fueia10jdra000e1c480/daots7ojdra000e1c490 +``` + +![lakeFS objects stored in the RustFS Console](./images/rustfs-lakefs-repo.png) + +## 5. Stop or reset the deployment + +Stop lakeFS while keeping the data: + +```bash +docker rm -f lakefs +``` + +The repository metadata and data stay in the `lakefs-data` bucket, so restarting lakeFS with the same configuration brings the repository back. To delete everything, remove the bucket: + +```bash +rc rb rustfs/lakefs-data --force +``` + +## Troubleshooting + +### `failed to create repository: failed to access storage` with a DNS error + +lakeFS is using virtual-hosted addressing. Set `force_path_style: true` under `blockstore.s3` — in the 1.x configuration schema the older `path_style` key is rejected at startup. + +### `missing required keys: [auth.encrypt.secret_key]` + +lakeFS 1.x requires an encryption key for the local database. Add the `auth.encrypt.secret_key` block as shown in the configuration above. + +### `mkdir /lakefs: permission denied` + +The container runs as a non-root user and cannot create the database directory. Run the container with `-u 0` for local testing, or mount a writable directory at the configured path. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional lakeFS operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [lakeFS S3 blockstore documentation](https://docs.lakefs.io/howto/using-s3.html) to configure the S3 gateway for tools that speak the S3 protocol. diff --git a/content/de/developer/integration/big-data/meta.json b/content/de/developer/integration/big-data/meta.json index ee7d69b5..249e219d 100644 --- a/content/de/developer/integration/big-data/meta.json +++ b/content/de/developer/integration/big-data/meta.json @@ -2,16 +2,18 @@ "title": "Data Analytics", "pages": [ "clickhouse", + "duckdb", + "doris", + "flink", + "hudi", "iceberg", - "pyiceberg", + "influxdb", + "lakefs", "milvus", "mlflow", "opendal", - "duckdb", - "doris", - "influxdb", + "pyiceberg", "spark", - "flink", "trino", "zeppelin" ] diff --git a/content/de/developer/integration/index.md b/content/de/developer/integration/index.md index e409e2e8..a7e68a04 100644 --- a/content/de/developer/integration/index.md +++ b/content/de/developer/integration/index.md @@ -8,8 +8,9 @@ Use this section to connect **RustFS** to infrastructure and application platfor ## Integration categories - [Reverse Proxy](./reverse-proxy/index.md) covers Nginx, Traefik, Caddy, and HAProxy. -- [Backup](./backup/index.md) covers Restic and Longhorn. -- [Datenanalyse](./big-data/index.md) covers analytics systems including ClickHouse, Doris, Iceberg, Milvus, OpenDAL, and Zeppelin. +- [Backup](./backup/index.md) covers Kopia, Longhorn, Restic, and Velero. +- [AI](./ai/index.md) covers AI platforms including Ray. +- [Datenanalyse](./big-data/index.md) covers analytics systems including ClickHouse, Doris, Hudi, Iceberg, lakeFS, Milvus, OpenDAL, and Zeppelin. - [Observability](./observability/index.md) covers telemetry systems including Fluentd, OpenObserve, OpenTelemetry, Thanos, and Tempo. - [Others](./others/index.md) covers the community-driven capo SDK for Python. - [Registry](./registry/index.md) covers Harbor. diff --git a/content/de/developer/integration/meta.json b/content/de/developer/integration/meta.json index b4b29480..a3c989a5 100644 --- a/content/de/developer/integration/meta.json +++ b/content/de/developer/integration/meta.json @@ -4,6 +4,7 @@ "reverse-proxy", "backup", "big-data", + "ai", "observability", "others", "registry", diff --git a/content/en/developer/integration/ai/images/rustfs-ray-data.png b/content/en/developer/integration/ai/images/rustfs-ray-data.png new file mode 100644 index 00000000..394f4499 Binary files /dev/null and b/content/en/developer/integration/ai/images/rustfs-ray-data.png differ diff --git a/content/en/developer/integration/ai/index.md b/content/en/developer/integration/ai/index.md new file mode 100644 index 00000000..06cc8316 --- /dev/null +++ b/content/en/developer/integration/ai/index.md @@ -0,0 +1,12 @@ +--- +title: "AI" +description: "Connect AI platforms to RustFS through S3-compatible object storage interfaces." +--- + +Use **RustFS** as the object storage layer for AI and machine learning platforms that support an S3-compatible endpoint. + +## Platforms + +- [Ray](./ray.md) + +Keep training datasets and checkpoints in dedicated buckets, and use credentials scoped to the required bucket operations. diff --git a/content/en/developer/integration/ai/meta.json b/content/en/developer/integration/ai/meta.json new file mode 100644 index 00000000..573fc520 --- /dev/null +++ b/content/en/developer/integration/ai/meta.json @@ -0,0 +1,6 @@ +{ + "title": "AI", + "pages": [ + "ray" + ] +} diff --git a/content/en/developer/integration/ai/ray.md b/content/en/developer/integration/ai/ray.md new file mode 100644 index 00000000..e9f71b94 --- /dev/null +++ b/content/en/developer/integration/ai/ray.md @@ -0,0 +1,103 @@ +--- +title: "Ray" +description: "Use Ray Data with RustFS as S3-compatible storage for dataset writes and reads." +--- + +This guide connects [Ray](https://github.com/ray-project/ray) — the distributed AI and Python compute framework — to **RustFS** through Ray Data's S3 filesystem support. You will run a Ray job inside the official image, write a dataset as Parquet to a RustFS bucket, read it back, and verify the objects. The workflow was verified with `rayproject/ray:2.44.0-py311` (Ray 2.44, pyarrow filesystem) and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Job["Ray job"] -->|"ray.data"| DS["Dataset"] + DS -->|"Parquet files"| RustFS["RustFS :9000"] +``` + +Ray Data reads and writes datasets through pyarrow's `S3FileSystem`. Passing an `S3FileSystem` configured for RustFS redirects every dataset operation — Parquet, CSV, JSON — to the bucket. + +## 1. Create the job file + +Create the script, replacing all connection placeholders: + +```python title="ray_s3.py" +import ray +ray.init(ignore_reinit_error=True) + +import pandas as pd +from pyarrow.fs import S3FileSystem + +fs = S3FileSystem( + endpoint_override="http://:9000", + access_key="", + secret_key="", + region="us-east-1", +) + +df = pd.DataFrame({"id": range(5), "value": [x * 1.5 for x in range(5)]}) +ds = ray.data.from_pandas(df) +ds.write_parquet("my-bucket/ray-demo/events/", filesystem=fs) + +back = ray.data.read_parquet("my-bucket/ray-demo/events/", filesystem=fs).take_all() +print("rows:", len(back)) +print("sample:", back[0]) +ray.shutdown() +``` + +`endpoint_override` takes the full endpoint URL including the scheme. pyarrow's `S3FileSystem` uses path-style requests for custom endpoints, so no extra flag is needed. The same filesystem object works for `write_csv`, `read_json`, and the other Ray Data methods. + +## 2. Run the job + +Run the script in the Ray image on the same Docker network as RustFS: + +```bash +docker run --rm --network oo-rustfs_default \ + -v "$PWD/ray_s3.py":/tmp/ray_s3.py \ + rayproject/ray:2.44.0-py311 python /tmp/ray_s3.py +``` + +```text +rows: 5 +sample: {'id': 0, 'value': 0.0} +``` + +## 3. Verify objects in RustFS + +List the dataset prefix: + +```bash +rc ls rustfs/my-bucket/ray-demo/ -r +``` + +Ray Data wrote the dataset as a Parquet block: + +```text +ray-demo/events/0_000000_000000.parquet +``` + +![Ray dataset files stored in the RustFS Console](./images/rustfs-ray-data.png) + +## 4. Stop or reset + +Ray Data holds no state of its own. To delete the demo dataset: + +```bash +rc rm rustfs/my-bucket/ray-demo/ --recursive --force +``` + +## Troubleshooting + +### `Unable to connect to endpoint` or timeouts + +Confirm `endpoint_override` includes the scheme and is reachable from the Ray container. Inside a Compose network the hostname is `rustfs`; from the host use `http://localhost:9000`. + +### `Access Denied` on write + +Confirm the access key and secret key are passed to `S3FileSystem` itself — Ray does not read the container's AWS environment variables through this code path. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Ray operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Ray Data documentation](https://docs.ray.io/en/latest/data/data.html) to chain transformations, training ingestion, and checkpointing on the same bucket. diff --git a/content/en/developer/integration/backup/images/rustfs-kopia-repo.png b/content/en/developer/integration/backup/images/rustfs-kopia-repo.png new file mode 100644 index 00000000..4241891c Binary files /dev/null and b/content/en/developer/integration/backup/images/rustfs-kopia-repo.png differ diff --git a/content/en/developer/integration/backup/images/rustfs-velero-backups.png b/content/en/developer/integration/backup/images/rustfs-velero-backups.png new file mode 100644 index 00000000..c3e9b169 Binary files /dev/null and b/content/en/developer/integration/backup/images/rustfs-velero-backups.png differ diff --git a/content/en/developer/integration/backup/index.md b/content/en/developer/integration/backup/index.md index 9a5b5114..b4208b6e 100644 --- a/content/en/developer/integration/backup/index.md +++ b/content/en/developer/integration/backup/index.md @@ -8,6 +8,8 @@ Use **RustFS** as the object storage backend for backup tools that store reposit ## Systems - [Restic](./restic.md) +- [Velero](./velero.md) +- [Kopia](./kopia.md) - [Longhorn](./longhorn.md) Keep backup jobs in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations. \ No newline at end of file diff --git a/content/en/developer/integration/backup/kopia.md b/content/en/developer/integration/backup/kopia.md new file mode 100644 index 00000000..fed21e5a --- /dev/null +++ b/content/en/developer/integration/backup/kopia.md @@ -0,0 +1,125 @@ +--- +title: "Kopia" +description: "Back up files to RustFS with Kopia's S3 repository backend." +--- + +This guide connects [Kopia](https://github.com/kopia/kopia) — the open-source backup and restore tool — to **RustFS** as an S3 repository. You will create a repository in a RustFS bucket, take a snapshot of a directory, restore it into an empty directory, and compare checksums. The workflow was verified with `kopia/kopia:0.18.1` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Source["Source files"] -->|snapshot| Kopia["Kopia"] + Kopia -->|"encrypted blocks"| RustFS["RustFS :9000"] + Kopia -->|restore| Restore["Restored files"] +``` + +Kopia stores the repository format files and deduplicated, encrypted content blocks in the bucket. Restores read the blocks back and reassemble the original files, so the checksum of every restored file must match the source. + +## 1. Create the repository + +Create the bucket first, then initialize the Kopia repository inside it. The endpoint is a bare `host:port` (no scheme); `--disable-tls` switches the client to plain HTTP: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/kopia-backups + +docker run --rm --network oo-rustfs_default kopia/kopia:0.18.1 repository create s3 \ + --bucket kopia-backups \ + --access-key \ + --secret-access-key \ + --endpoint :9000 \ + --region us-east-1 \ + --disable-tls \ + --password \ + --override-username demo --override-hostname workstation +``` + +Kopia validates the provider by reading and writing through the S3 API before it reports success. + +## 2. Connect, snapshot, and restore + +Run the following commands from the directory holding your `repository.config` (created by the previous step). Kopia reads the connection settings from that file, so the S3 flags are only needed once: + +```bash +export KOPIA_PASSWORD= +export KOPIA_CONFIG_PATH=/config/repository.config + +alias kopia='docker run --rm --network oo-rustfs_default \ + -e KOPIA_PASSWORD -e KOPIA_CONFIG_PATH \ + -v "$PWD/config:/config" \ + -v "$PWD/source:/source:ro" \ + -v "$PWD/restore:/restore" kopia/kopia:0.18.1' + +kopia repository connect s3 \ + --bucket kopia-backups \ + --access-key \ + --secret-access-key \ + --endpoint :9000 \ + --region us-east-1 --disable-tls \ + --override-username demo --override-hostname workstation + +kopia snapshot create /source +kopia snapshot list +``` + +Restore the snapshot into an empty directory and compare checksums with the source. The snapshot ID is the `ka...` identifier printed by `snapshot list`: + +```bash +kopia restore /restore + +sha256sum source/blob.bin restore/blob.bin +``` + +```text +bec4530e2798465b... source/blob.bin +bec4530e2798465b... restore/blob.bin +``` + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/kopia-backups/ -r +``` + +The output shows the repository format files plus the packed content blocks written by the snapshot: + +```text +kopia.blobcfg +kopia.repository +p0000.../... +``` + +![Kopia repository blocks stored in the RustFS Console](./images/rustfs-kopia-repo.png) + +## 4. Stop or reset + +Kopia is a client-side tool and holds no running state. To delete the repository and all snapshots, remove the bucket: + +```bash +rc rb rustfs/kopia-backups --force +``` + +## Troubleshooting + +### `Endpoint url cannot have fully qualified paths` + +The endpoint must be a bare `host:port` value without a scheme or path — Kopia builds the object URLs itself. + +### `server gave HTTP response to HTTPS client` + +Without `--disable-tls`, Kopia speaks HTTPS. RustFS without TLS needs the `--disable-tls` flag on both `repository create s3` and `repository connect s3`. + +### `can't connect to storage` with a DNS error + +The S3 client is using virtual-hosted addressing. Keep the endpoint as a bare `host:port` value; Kopia uses path-style requests for endpoints given in that form. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Kopia operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Kopia repository documentation](https://kopia.io/docs/repositories/) to add policies, retention, and scheduled snapshots. diff --git a/content/en/developer/integration/backup/meta.json b/content/en/developer/integration/backup/meta.json index 3ea20cbe..dac5156c 100644 --- a/content/en/developer/integration/backup/meta.json +++ b/content/en/developer/integration/backup/meta.json @@ -1,7 +1,9 @@ { "title": "Backup", "pages": [ + "kopia", + "longhorn", "restic", - "longhorn" + "velero" ] } diff --git a/content/en/developer/integration/backup/velero.md b/content/en/developer/integration/backup/velero.md new file mode 100644 index 00000000..0140296b --- /dev/null +++ b/content/en/developer/integration/backup/velero.md @@ -0,0 +1,153 @@ +--- +title: "Velero" +description: "Back up Kubernetes cluster resources to RustFS with Velero's AWS object store provider." +--- + +This guide connects [Velero](https://github.com/vmware-tanzu/velero) — the Kubernetes backup and restore tool — to **RustFS** as its object storage backend. You will install Velero into a Kubernetes cluster with a RustFS backup location, back up cluster resources, delete them, restore from the bucket, and verify the restored objects. The workflow was verified with Velero CLI 1.16.2, `velero-plugin-for-aws:v1.12.2` on k3s (Kubernetes 1.30), and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need a Kubernetes cluster with `kubectl` access and the Velero CLI. This guide is intended for integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + K8s["Kubernetes cluster"] -->|"resources"| V["Velero"] + V -->|"backups + logs"| RustFS["RustFS :9000"] +``` + +Velero serializes Kubernetes resources and (optionally) pod volume data into gzip archives under `backups//` in the bucket. Restores download those archives and recreate the resources in the cluster. + +## 1. Create the backup bucket + +Create a dedicated bucket — Velero does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/velero-backups +``` + +## 2. Install Velero + +Write the credentials to a file and install the Velero server. The `s3Url` must be an endpoint reachable from the cluster pods themselves — use the node IP or an internal address, not a port forward from your workstation: + +```bash +cat > velero-creds < +aws_secret_access_key= +EOF + +kubectl create namespace velero +kubectl create secret generic cloud-credentials \ + --namespace velero --from-file=cloud=velero-creds + +velero install \ + --provider aws \ + --plugins velero/velero-plugin-for-aws:v1.12.2 \ + --bucket velero-backups \ + --backup-location-config region=us-east-1,s3ForcePathStyle="true",s3Url=http://:9000 \ + --secret-file velero-creds \ + --use-volume-snapshots=false +``` + +Wait for the deployment to become ready and the backup location to turn `Available`: + +```bash +kubectl -n velero get pods +velero backup-location get +``` + +```text +NAME PROVIDER BUCKET/PREFIX PHASE LAST VALIDATED ACCESS MODE DEFAULT +default aws velero-backups Available 2026-09-22 ... ReadWrite true +``` + +## 3. Back up cluster resources + +Create two demo resources and back up the whole default namespace scope: + +```bash +kubectl create configmap demo-cm --from-literal=key=rustfs-velero-demo +kubectl create deployment nginx --image=nginx:1.27 + +velero backup create demo-backup --wait +velero backup get +``` + +The backup uploads the resource archives and logs to RustFS. + +## 4. Verify the backup in RustFS + +List the bucket: + +```bash +rc ls rustfs/velero-backups/ -r +``` + +The backup archive set is stored under the `backups/` prefix: + +```text +backups/demo-backup/demo-backup-resources.json.gz +backups/demo-backup/demo-backup-logs.gz +backups/demo-backup/demo-backup-itemoperations.json.gz +``` + +![Velero backups stored in the RustFS Console](./images/rustfs-velero-backups.png) + +## 5. Restore and verify + +Delete the demo resources and restore them from the backup. The restore reads the archives from RustFS and recreates the resources: + +```bash +kubectl delete configmap demo-cm +kubectl delete deployment nginx + +velero restore create --from-backup demo-backup --wait +``` + +Confirm the resources are back with their original content: + +```bash +kubectl get configmap demo-cm -o jsonpath="{.data.key}" +kubectl get deployment nginx +``` + +```text +rustfs-velero-demo +NAME READY UP-TO-DATE AVAILABLE AGE +nginx 1/1 1 1 10s +``` + +## 6. Stop or reset + +The backups stay in the `velero-backups` bucket and restore into any cluster that runs Velero against the same bucket. To remove the demo backup from RustFS: + +```bash +velero backup delete demo-backup --confirm +rc rm rustfs/velero-backups/backups/ --recursive --force +``` + +## Troubleshooting + +### The backup location stays `Unavailable` + +The Velero pod cannot reach the `s3Url`. Endpoints on a Docker custom bridge network are not routable from Kubernetes pods — use the Docker host bridge IP (for example `http://172.17.0.1:9000` on a single-host cluster) or a node address. Patch the location and let Velero revalidate: + +```bash +kubectl -n velero patch backupstoragelocation default --type merge -p \ + '{"spec":{"config":{"s3Url":"http://172.17.0.1:9000"}}}' +``` + +### The backup ends `PartiallyFailed` with a daemonset error + +The error `daemonset pod not found in running state` means the node-agent pod (used for pod volume backups) was not running. It does not affect the cluster-resource archives. Wait for the agent or keep `--default-volumes-to-fs-backup` only when the agent is running. + +### `FailedValidation` right after install + +The first backup attempt runs against a location that has not validated yet. Wait for `velero backup-location get` to report `Available` before creating backups. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Velero operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Velero documentation](https://velero.io/docs/main/) to add schedules, volume snapshots, and cluster migration flows. diff --git a/content/en/developer/integration/big-data/hudi.md b/content/en/developer/integration/big-data/hudi.md new file mode 100644 index 00000000..42026454 --- /dev/null +++ b/content/en/developer/integration/big-data/hudi.md @@ -0,0 +1,109 @@ +--- +title: "Apache Hudi" +description: "Write Apache Hudi tables to RustFS through Spark and the s3a connector." +--- + +This guide connects [Apache Hudi](https://github.com/apache/hudi) — the transactional data lake platform — to **RustFS** as its copy-on-write table store. You will run Spark with the Hudi bundle, write a table to the `s3a://` location inside a RustFS bucket, read it back, and verify the table files. The workflow was verified with `apache/spark:3.5.6`, `hudi-spark3.5-bundle_2.12:0.15.0`, `hadoop-aws:3.3.4`, and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Job["Spark + Hudi"] -->|"commit + parquet"| RustFS["RustFS :9000"] +``` + +Hudi stores the `.hoodie/` timeline, commit files, and Parquet data blocks under the table path in the bucket. Reads resolve the latest table snapshot from the timeline, so every write and query goes through the S3 API. + +## 1. Run the Spark write + +Hudi requires the Kryo serializer and the s3a credentials as Hadoop properties. The following `spark-shell` session creates a non-partitioned table in `my-bucket` — replace all connection placeholders: + +```scala +import org.apache.spark.sql.SaveMode + +val df = Seq((1, "a"), (2, "b"), (3, "c")).toDF("id", "name") +df.write.format("hudi") + .option("hoodie.table.name", "events") + .option("hoodie.datasource.write.recordkey.field", "id") + .option("hoodie.datasource.write.precombine.field", "name") + .option("hoodie.datasource.write.partitionpath.field", "") + .mode(SaveMode.Overwrite) + .save("s3a://my-bucket/hudi-demo/events") + +val back = spark.read.format("hudi").load("s3a://my-bucket/hudi-demo/events") +back.select("id", "name").show() +``` + +```bash +docker run --rm --network oo-rustfs_default \ + -v "$PWD/hudi_test.scala":/tmp/hudi_test.scala \ + apache/spark:3.5.6 /opt/spark/bin/spark-shell \ + --packages org.apache.hudi:hudi-spark3.5-bundle_2.12:0.15.0,org.apache.hadoop:hadoop-aws:3.3.4 \ + --conf spark.serializer=org.apache.spark.serializer.KryoSerializer \ + --conf spark.hadoop.fs.s3a.endpoint=http://rustfs:9000 \ + --conf spark.hadoop.fs.s3a.path.style.access=true \ + --conf spark.hadoop.fs.s3a.access.key= \ + --conf spark.hadoop.fs.s3a.secret.key= \ + --conf spark.hadoop.fs.s3a.region=us-east-1 \ + --conf spark.jars.ivy=/tmp/.ivy2 \ + --conf spark.sql.shuffle.partitions=2 \ + -i /tmp/hudi_test.scala +``` + +The `--packages` flags download the Hudi bundle and the s3a connector on first run. The `spark.jars.ivy` setting avoids Ivy cache permission errors in the container. The read-back `show()` prints the three rows: + +```text ++---+----+ +| 2| b| +| 3| c| +| 1| a| ++---+----+ +``` + +## 2. Verify objects in RustFS + +List the table prefix: + +```bash +rc ls rustfs/my-bucket/hudi-demo/ -r +``` + +The Hudi table layout appears under the table path — the `.hoodie/` timeline with the commit file, plus the Parquet data block: + +```text +hudi-demo/events/.hoodie/20260922015803291.commit +hudi-demo/events/.hoodie/20260922015803291.commit.requested +hudi-demo/events/// +``` + +![Hudi table files stored in the RustFS Console](./images/rustfs-hudi-table.png) + +## 3. Stop or reset + +Spark runs as a one-shot client and holds no state. To delete the demo table: + +```bash +rc rm rustfs/my-bucket/hudi-demo/ --recursive --force +``` + +## Troubleshooting + +### `hoodie only support org.apache.spark.serializer.KryoSerializer as spark.serializer` + +Hudi rejects the default Spark serializer. Pass `--conf spark.serializer=org.apache.spark.serializer.KryoSerializer` as shown above. + +### `Partition-path field has to be non-empty` or keygenerator class errors + +For a non-partitioned table, keep the default key generator (do not set `hoodie.datasource.write.keygenerator.class`) and set `hoodie.datasource.write.partitionpath.field` to the empty string. + +### `NoSuchMethodError` or `ClassNotFoundException` in the Hudi write + +The Spark and Hudi bundle versions must match: `hudi-spark3.5-bundle_2.12` goes with Spark 3.5.x, and the bundled Hadoop client must be compatible with `hadoop-aws:3.3.4`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Hudi operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Hudi Spark guide](https://hudi.apache.org/docs/quick-start-guide/) to add upserts, compaction, and query integrations on the same table. diff --git a/content/en/developer/integration/big-data/images/rustfs-hudi-table.png b/content/en/developer/integration/big-data/images/rustfs-hudi-table.png new file mode 100644 index 00000000..b91c9157 Binary files /dev/null and b/content/en/developer/integration/big-data/images/rustfs-hudi-table.png differ diff --git a/content/en/developer/integration/big-data/images/rustfs-lakefs-repo.png b/content/en/developer/integration/big-data/images/rustfs-lakefs-repo.png new file mode 100644 index 00000000..0236a479 Binary files /dev/null and b/content/en/developer/integration/big-data/images/rustfs-lakefs-repo.png differ diff --git a/content/en/developer/integration/big-data/index.md b/content/en/developer/integration/big-data/index.md index bb96ab87..ba66ba0d 100644 --- a/content/en/developer/integration/big-data/index.md +++ b/content/en/developer/integration/big-data/index.md @@ -8,6 +8,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo ## Systems - [ClickHouse](./clickhouse.md) +- [Hudi](./hudi.md) - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) - [Milvus](./milvus.md) @@ -15,6 +16,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [OpenDAL](./opendal.md) - [DuckDB](./duckdb.md) - [Doris](./doris.md) +- [lakeFS](./lakefs.md) - [InfluxDB](./influxdb.md) - [Spark](./spark.md) - [Flink](./flink.md) diff --git a/content/en/developer/integration/big-data/lakefs.md b/content/en/developer/integration/big-data/lakefs.md new file mode 100644 index 00000000..ba63ef22 --- /dev/null +++ b/content/en/developer/integration/big-data/lakefs.md @@ -0,0 +1,160 @@ +--- +title: "lakeFS" +description: "Run lakeFS with RustFS as its S3 blockstore for versioned data lakes." +--- + +This guide connects [lakeFS](https://github.com/treeverse/lakeFS) — the Git-like data lake versioning layer — to **RustFS** as its S3 blockstore. You will start lakeFS, create a repository whose storage namespace points at a RustFS bucket, commit an object, and verify that the lakeFS metadata and data files live in RustFS. The workflow was verified with `treeverse/lakefs:1.58.0` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["lakectl / API"] -->|HTTP| LakeFS["lakeFS :8000"] + LakeFS -->|"metadata + data"| RustFS["RustFS :9000"] +``` + +lakeFS stores repository metadata and committed data files in the blockstore under the storage namespace. The bucket holds one `repo/` prefix containing `_lakefs/` metadata and `data/` objects; lakeFS reads and writes them through the S3 API. + +## 1. Create the project files + +Create the bucket, then the lakeFS configuration: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/lakefs-data +``` + +```yaml title="config.yaml" +database: + type: local + local: + path: /lakefs/data + +blockstore: + type: s3 + s3: + endpoint: http://rustfs:9000 + region: us-east-1 + force_path_style: true + +auth: + encrypt: + secret_key: change-me-to-a-random-string + +gateways: + s3: + domain_name: rustfs-gateway.local + +logging: + format: text + level: info +``` + +`force_path_style: true` is required — without it lakeFS builds `lakefs-data.` as a hostname and every request fails with a DNS error. The credentials are supplied through the standard `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` environment variables. + +Create the environment file for the service credentials and the initial admin user, replacing the placeholders: + +```ini title=".env" +AWS_ACCESS_KEY_ID= +AWS_SECRET_ACCESS_KEY= +LAKEFS_INSTALLATION_USER_NAME=admin +LAKEFS_INSTALLATION_ACCESS_KEY_ID= +LAKEFS_INSTALLATION_SECRET_ACCESS_KEY= +LAKEFS_STATS_ENABLED=false +``` + +## 2. Start lakeFS + +```bash +docker run -d --name lakefs --network oo-rustfs_default \ + -p 8000:8000 \ + -v "$PWD/config.yaml":/etc/lakefs/config.yaml:ro \ + --env-file .env \ + treeverse/lakefs:1.58.0 run +``` + +Wait for `http://localhost:8000/api/healthcheck` to return `200`, then create a repository with a storage namespace inside the bucket: + +```bash +curl -s -u : \ + -X POST http://localhost:8000/api/v1/repositories \ + -H "Content-Type: application/json" \ + -d '{"name": "rustfs-demo", "storage_namespace": "s3://lakefs-data/repo"}' +``` + +## 3. Commit an object + +Upload a file to the `main` branch and commit it: + +```bash +echo "hello from lakefs on rustfs" > hello.txt + +curl -s -u : \ + -X POST "http://localhost:8000/api/v1/repositories/rustfs-demo/branches/main/objects?path=hello.txt" \ + --data-binary @hello.txt + +curl -s -u : \ + -X POST "http://localhost:8000/api/v1/repositories/rustfs-demo/branches/main/commits" \ + -H "Content-Type: application/json" \ + -d '{"message": "add hello"}' +``` + +Read the object back through the branch ref — the response body is the committed content: + +```bash +curl -s -u : \ + "http://localhost:8000/api/v1/repositories/rustfs-demo/refs/main/objects?path=hello.txt" +``` + +## 4. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/lakefs-data/ -r +``` + +The output shows the lakeFS metadata objects and the committed data file under the repository prefix: + +```text +repo/_lakefs/19b2b26e37cb20fc6763c527f88eb5151891b04a2c8c9ddd32870c5c3f353281 +repo/data/fueia10jdra000e1c480/daots7ojdra000e1c490 +``` + +![lakeFS objects stored in the RustFS Console](./images/rustfs-lakefs-repo.png) + +## 5. Stop or reset the deployment + +Stop lakeFS while keeping the data: + +```bash +docker rm -f lakefs +``` + +The repository metadata and data stay in the `lakefs-data` bucket, so restarting lakeFS with the same configuration brings the repository back. To delete everything, remove the bucket: + +```bash +rc rb rustfs/lakefs-data --force +``` + +## Troubleshooting + +### `failed to create repository: failed to access storage` with a DNS error + +lakeFS is using virtual-hosted addressing. Set `force_path_style: true` under `blockstore.s3` — in the 1.x configuration schema the older `path_style` key is rejected at startup. + +### `missing required keys: [auth.encrypt.secret_key]` + +lakeFS 1.x requires an encryption key for the local database. Add the `auth.encrypt.secret_key` block as shown in the configuration above. + +### `mkdir /lakefs: permission denied` + +The container runs as a non-root user and cannot create the database directory. Run the container with `-u 0` for local testing, or mount a writable directory at the configured path. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional lakeFS operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [lakeFS S3 blockstore documentation](https://docs.lakefs.io/howto/using-s3.html) to configure the S3 gateway for tools that speak the S3 protocol. diff --git a/content/en/developer/integration/big-data/meta.json b/content/en/developer/integration/big-data/meta.json index ee7d69b5..249e219d 100644 --- a/content/en/developer/integration/big-data/meta.json +++ b/content/en/developer/integration/big-data/meta.json @@ -2,16 +2,18 @@ "title": "Data Analytics", "pages": [ "clickhouse", + "duckdb", + "doris", + "flink", + "hudi", "iceberg", - "pyiceberg", + "influxdb", + "lakefs", "milvus", "mlflow", "opendal", - "duckdb", - "doris", - "influxdb", + "pyiceberg", "spark", - "flink", "trino", "zeppelin" ] diff --git a/content/en/developer/integration/index.md b/content/en/developer/integration/index.md index 38f2ac3d..92980c37 100644 --- a/content/en/developer/integration/index.md +++ b/content/en/developer/integration/index.md @@ -8,8 +8,9 @@ Use this section to connect **RustFS** to infrastructure and application platfor ## Integration categories - [Reverse Proxy](./reverse-proxy/index.md) covers Nginx, Traefik, Caddy, and HAProxy. -- [Backup](./backup/index.md) covers Restic and Longhorn. -- [Data Analytics](./big-data/index.md) covers analytics systems including ClickHouse, Doris, Iceberg, Milvus, OpenDAL, and Zeppelin. +- [Backup](./backup/index.md) covers Kopia, Longhorn, Restic, and Velero. +- [AI](./ai/index.md) covers AI platforms including Ray. +- [Data Analytics](./big-data/index.md) covers analytics systems including ClickHouse, Doris, Hudi, Iceberg, lakeFS, Milvus, OpenDAL, and Zeppelin. - [Observability](./observability/index.md) covers telemetry systems including Fluentd, OpenObserve, OpenTelemetry, Thanos, and Tempo. - [Others](./others/index.md) covers the community-driven capo SDK for Python. - [Registry](./registry/index.md) covers Harbor. diff --git a/content/en/developer/integration/meta.json b/content/en/developer/integration/meta.json index b4b29480..a3c989a5 100644 --- a/content/en/developer/integration/meta.json +++ b/content/en/developer/integration/meta.json @@ -4,6 +4,7 @@ "reverse-proxy", "backup", "big-data", + "ai", "observability", "others", "registry", diff --git a/content/fr/developer/integration/ai/images/rustfs-ray-data.png b/content/fr/developer/integration/ai/images/rustfs-ray-data.png new file mode 100644 index 00000000..394f4499 Binary files /dev/null and b/content/fr/developer/integration/ai/images/rustfs-ray-data.png differ diff --git a/content/fr/developer/integration/ai/index.md b/content/fr/developer/integration/ai/index.md new file mode 100644 index 00000000..37fa7892 --- /dev/null +++ b/content/fr/developer/integration/ai/index.md @@ -0,0 +1,12 @@ +--- +title: "IA" +description: "Connectez les plateformes d'IA à RustFS via des interfaces de stockage objet compatibles S3." +--- + +Utilisez **RustFS** comme couche de stockage objet pour les plateformes d'IA et de machine learning qui prennent en charge un point de terminaison compatible S3. + +## Plateformes + +- [Ray](./ray.md) + +Conservez les jeux de données et les checkpoints dans des buckets dédiés et limitez les identifiants aux opérations de bucket requises. diff --git a/content/fr/developer/integration/ai/meta.json b/content/fr/developer/integration/ai/meta.json new file mode 100644 index 00000000..876f2b47 --- /dev/null +++ b/content/fr/developer/integration/ai/meta.json @@ -0,0 +1,6 @@ +{ + "title": "IA", + "pages": [ + "ray" + ] +} diff --git a/content/fr/developer/integration/ai/ray.md b/content/fr/developer/integration/ai/ray.md new file mode 100644 index 00000000..e9f71b94 --- /dev/null +++ b/content/fr/developer/integration/ai/ray.md @@ -0,0 +1,103 @@ +--- +title: "Ray" +description: "Use Ray Data with RustFS as S3-compatible storage for dataset writes and reads." +--- + +This guide connects [Ray](https://github.com/ray-project/ray) — the distributed AI and Python compute framework — to **RustFS** through Ray Data's S3 filesystem support. You will run a Ray job inside the official image, write a dataset as Parquet to a RustFS bucket, read it back, and verify the objects. The workflow was verified with `rayproject/ray:2.44.0-py311` (Ray 2.44, pyarrow filesystem) and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Job["Ray job"] -->|"ray.data"| DS["Dataset"] + DS -->|"Parquet files"| RustFS["RustFS :9000"] +``` + +Ray Data reads and writes datasets through pyarrow's `S3FileSystem`. Passing an `S3FileSystem` configured for RustFS redirects every dataset operation — Parquet, CSV, JSON — to the bucket. + +## 1. Create the job file + +Create the script, replacing all connection placeholders: + +```python title="ray_s3.py" +import ray +ray.init(ignore_reinit_error=True) + +import pandas as pd +from pyarrow.fs import S3FileSystem + +fs = S3FileSystem( + endpoint_override="http://:9000", + access_key="", + secret_key="", + region="us-east-1", +) + +df = pd.DataFrame({"id": range(5), "value": [x * 1.5 for x in range(5)]}) +ds = ray.data.from_pandas(df) +ds.write_parquet("my-bucket/ray-demo/events/", filesystem=fs) + +back = ray.data.read_parquet("my-bucket/ray-demo/events/", filesystem=fs).take_all() +print("rows:", len(back)) +print("sample:", back[0]) +ray.shutdown() +``` + +`endpoint_override` takes the full endpoint URL including the scheme. pyarrow's `S3FileSystem` uses path-style requests for custom endpoints, so no extra flag is needed. The same filesystem object works for `write_csv`, `read_json`, and the other Ray Data methods. + +## 2. Run the job + +Run the script in the Ray image on the same Docker network as RustFS: + +```bash +docker run --rm --network oo-rustfs_default \ + -v "$PWD/ray_s3.py":/tmp/ray_s3.py \ + rayproject/ray:2.44.0-py311 python /tmp/ray_s3.py +``` + +```text +rows: 5 +sample: {'id': 0, 'value': 0.0} +``` + +## 3. Verify objects in RustFS + +List the dataset prefix: + +```bash +rc ls rustfs/my-bucket/ray-demo/ -r +``` + +Ray Data wrote the dataset as a Parquet block: + +```text +ray-demo/events/0_000000_000000.parquet +``` + +![Ray dataset files stored in the RustFS Console](./images/rustfs-ray-data.png) + +## 4. Stop or reset + +Ray Data holds no state of its own. To delete the demo dataset: + +```bash +rc rm rustfs/my-bucket/ray-demo/ --recursive --force +``` + +## Troubleshooting + +### `Unable to connect to endpoint` or timeouts + +Confirm `endpoint_override` includes the scheme and is reachable from the Ray container. Inside a Compose network the hostname is `rustfs`; from the host use `http://localhost:9000`. + +### `Access Denied` on write + +Confirm the access key and secret key are passed to `S3FileSystem` itself — Ray does not read the container's AWS environment variables through this code path. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Ray operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Ray Data documentation](https://docs.ray.io/en/latest/data/data.html) to chain transformations, training ingestion, and checkpointing on the same bucket. diff --git a/content/fr/developer/integration/backup/images/rustfs-kopia-repo.png b/content/fr/developer/integration/backup/images/rustfs-kopia-repo.png new file mode 100644 index 00000000..4241891c Binary files /dev/null and b/content/fr/developer/integration/backup/images/rustfs-kopia-repo.png differ diff --git a/content/fr/developer/integration/backup/images/rustfs-velero-backups.png b/content/fr/developer/integration/backup/images/rustfs-velero-backups.png new file mode 100644 index 00000000..c3e9b169 Binary files /dev/null and b/content/fr/developer/integration/backup/images/rustfs-velero-backups.png differ diff --git a/content/fr/developer/integration/backup/index.md b/content/fr/developer/integration/backup/index.md index b29cbf60..6b837c02 100644 --- a/content/fr/developer/integration/backup/index.md +++ b/content/fr/developer/integration/backup/index.md @@ -8,6 +8,8 @@ Utilisez **RustFS** comme backend de stockage objet pour les outils de sauvegard ## Systèmes - [Restic](./restic.md) +- [Velero](./velero.md) +- [Kopia](./kopia.md) - [Longhorn](./longhorn.md) Conservez les tâches de sauvegarde dans un compartiment et un préfixe dédiés, et utilisez des identifiants limités aux opérations de compartiment nécessaires. \ No newline at end of file diff --git a/content/fr/developer/integration/backup/kopia.md b/content/fr/developer/integration/backup/kopia.md new file mode 100644 index 00000000..fed21e5a --- /dev/null +++ b/content/fr/developer/integration/backup/kopia.md @@ -0,0 +1,125 @@ +--- +title: "Kopia" +description: "Back up files to RustFS with Kopia's S3 repository backend." +--- + +This guide connects [Kopia](https://github.com/kopia/kopia) — the open-source backup and restore tool — to **RustFS** as an S3 repository. You will create a repository in a RustFS bucket, take a snapshot of a directory, restore it into an empty directory, and compare checksums. The workflow was verified with `kopia/kopia:0.18.1` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Source["Source files"] -->|snapshot| Kopia["Kopia"] + Kopia -->|"encrypted blocks"| RustFS["RustFS :9000"] + Kopia -->|restore| Restore["Restored files"] +``` + +Kopia stores the repository format files and deduplicated, encrypted content blocks in the bucket. Restores read the blocks back and reassemble the original files, so the checksum of every restored file must match the source. + +## 1. Create the repository + +Create the bucket first, then initialize the Kopia repository inside it. The endpoint is a bare `host:port` (no scheme); `--disable-tls` switches the client to plain HTTP: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/kopia-backups + +docker run --rm --network oo-rustfs_default kopia/kopia:0.18.1 repository create s3 \ + --bucket kopia-backups \ + --access-key \ + --secret-access-key \ + --endpoint :9000 \ + --region us-east-1 \ + --disable-tls \ + --password \ + --override-username demo --override-hostname workstation +``` + +Kopia validates the provider by reading and writing through the S3 API before it reports success. + +## 2. Connect, snapshot, and restore + +Run the following commands from the directory holding your `repository.config` (created by the previous step). Kopia reads the connection settings from that file, so the S3 flags are only needed once: + +```bash +export KOPIA_PASSWORD= +export KOPIA_CONFIG_PATH=/config/repository.config + +alias kopia='docker run --rm --network oo-rustfs_default \ + -e KOPIA_PASSWORD -e KOPIA_CONFIG_PATH \ + -v "$PWD/config:/config" \ + -v "$PWD/source:/source:ro" \ + -v "$PWD/restore:/restore" kopia/kopia:0.18.1' + +kopia repository connect s3 \ + --bucket kopia-backups \ + --access-key \ + --secret-access-key \ + --endpoint :9000 \ + --region us-east-1 --disable-tls \ + --override-username demo --override-hostname workstation + +kopia snapshot create /source +kopia snapshot list +``` + +Restore the snapshot into an empty directory and compare checksums with the source. The snapshot ID is the `ka...` identifier printed by `snapshot list`: + +```bash +kopia restore /restore + +sha256sum source/blob.bin restore/blob.bin +``` + +```text +bec4530e2798465b... source/blob.bin +bec4530e2798465b... restore/blob.bin +``` + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/kopia-backups/ -r +``` + +The output shows the repository format files plus the packed content blocks written by the snapshot: + +```text +kopia.blobcfg +kopia.repository +p0000.../... +``` + +![Kopia repository blocks stored in the RustFS Console](./images/rustfs-kopia-repo.png) + +## 4. Stop or reset + +Kopia is a client-side tool and holds no running state. To delete the repository and all snapshots, remove the bucket: + +```bash +rc rb rustfs/kopia-backups --force +``` + +## Troubleshooting + +### `Endpoint url cannot have fully qualified paths` + +The endpoint must be a bare `host:port` value without a scheme or path — Kopia builds the object URLs itself. + +### `server gave HTTP response to HTTPS client` + +Without `--disable-tls`, Kopia speaks HTTPS. RustFS without TLS needs the `--disable-tls` flag on both `repository create s3` and `repository connect s3`. + +### `can't connect to storage` with a DNS error + +The S3 client is using virtual-hosted addressing. Keep the endpoint as a bare `host:port` value; Kopia uses path-style requests for endpoints given in that form. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Kopia operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Kopia repository documentation](https://kopia.io/docs/repositories/) to add policies, retention, and scheduled snapshots. diff --git a/content/fr/developer/integration/backup/meta.json b/content/fr/developer/integration/backup/meta.json index b9923706..6634f836 100644 --- a/content/fr/developer/integration/backup/meta.json +++ b/content/fr/developer/integration/backup/meta.json @@ -1,7 +1,9 @@ { "title": "Sauvegarde", "pages": [ + "kopia", + "longhorn", "restic", - "longhorn" + "velero" ] } diff --git a/content/fr/developer/integration/backup/velero.md b/content/fr/developer/integration/backup/velero.md new file mode 100644 index 00000000..0140296b --- /dev/null +++ b/content/fr/developer/integration/backup/velero.md @@ -0,0 +1,153 @@ +--- +title: "Velero" +description: "Back up Kubernetes cluster resources to RustFS with Velero's AWS object store provider." +--- + +This guide connects [Velero](https://github.com/vmware-tanzu/velero) — the Kubernetes backup and restore tool — to **RustFS** as its object storage backend. You will install Velero into a Kubernetes cluster with a RustFS backup location, back up cluster resources, delete them, restore from the bucket, and verify the restored objects. The workflow was verified with Velero CLI 1.16.2, `velero-plugin-for-aws:v1.12.2` on k3s (Kubernetes 1.30), and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need a Kubernetes cluster with `kubectl` access and the Velero CLI. This guide is intended for integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + K8s["Kubernetes cluster"] -->|"resources"| V["Velero"] + V -->|"backups + logs"| RustFS["RustFS :9000"] +``` + +Velero serializes Kubernetes resources and (optionally) pod volume data into gzip archives under `backups//` in the bucket. Restores download those archives and recreate the resources in the cluster. + +## 1. Create the backup bucket + +Create a dedicated bucket — Velero does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/velero-backups +``` + +## 2. Install Velero + +Write the credentials to a file and install the Velero server. The `s3Url` must be an endpoint reachable from the cluster pods themselves — use the node IP or an internal address, not a port forward from your workstation: + +```bash +cat > velero-creds < +aws_secret_access_key= +EOF + +kubectl create namespace velero +kubectl create secret generic cloud-credentials \ + --namespace velero --from-file=cloud=velero-creds + +velero install \ + --provider aws \ + --plugins velero/velero-plugin-for-aws:v1.12.2 \ + --bucket velero-backups \ + --backup-location-config region=us-east-1,s3ForcePathStyle="true",s3Url=http://:9000 \ + --secret-file velero-creds \ + --use-volume-snapshots=false +``` + +Wait for the deployment to become ready and the backup location to turn `Available`: + +```bash +kubectl -n velero get pods +velero backup-location get +``` + +```text +NAME PROVIDER BUCKET/PREFIX PHASE LAST VALIDATED ACCESS MODE DEFAULT +default aws velero-backups Available 2026-09-22 ... ReadWrite true +``` + +## 3. Back up cluster resources + +Create two demo resources and back up the whole default namespace scope: + +```bash +kubectl create configmap demo-cm --from-literal=key=rustfs-velero-demo +kubectl create deployment nginx --image=nginx:1.27 + +velero backup create demo-backup --wait +velero backup get +``` + +The backup uploads the resource archives and logs to RustFS. + +## 4. Verify the backup in RustFS + +List the bucket: + +```bash +rc ls rustfs/velero-backups/ -r +``` + +The backup archive set is stored under the `backups/` prefix: + +```text +backups/demo-backup/demo-backup-resources.json.gz +backups/demo-backup/demo-backup-logs.gz +backups/demo-backup/demo-backup-itemoperations.json.gz +``` + +![Velero backups stored in the RustFS Console](./images/rustfs-velero-backups.png) + +## 5. Restore and verify + +Delete the demo resources and restore them from the backup. The restore reads the archives from RustFS and recreates the resources: + +```bash +kubectl delete configmap demo-cm +kubectl delete deployment nginx + +velero restore create --from-backup demo-backup --wait +``` + +Confirm the resources are back with their original content: + +```bash +kubectl get configmap demo-cm -o jsonpath="{.data.key}" +kubectl get deployment nginx +``` + +```text +rustfs-velero-demo +NAME READY UP-TO-DATE AVAILABLE AGE +nginx 1/1 1 1 10s +``` + +## 6. Stop or reset + +The backups stay in the `velero-backups` bucket and restore into any cluster that runs Velero against the same bucket. To remove the demo backup from RustFS: + +```bash +velero backup delete demo-backup --confirm +rc rm rustfs/velero-backups/backups/ --recursive --force +``` + +## Troubleshooting + +### The backup location stays `Unavailable` + +The Velero pod cannot reach the `s3Url`. Endpoints on a Docker custom bridge network are not routable from Kubernetes pods — use the Docker host bridge IP (for example `http://172.17.0.1:9000` on a single-host cluster) or a node address. Patch the location and let Velero revalidate: + +```bash +kubectl -n velero patch backupstoragelocation default --type merge -p \ + '{"spec":{"config":{"s3Url":"http://172.17.0.1:9000"}}}' +``` + +### The backup ends `PartiallyFailed` with a daemonset error + +The error `daemonset pod not found in running state` means the node-agent pod (used for pod volume backups) was not running. It does not affect the cluster-resource archives. Wait for the agent or keep `--default-volumes-to-fs-backup` only when the agent is running. + +### `FailedValidation` right after install + +The first backup attempt runs against a location that has not validated yet. Wait for `velero backup-location get` to report `Available` before creating backups. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Velero operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Velero documentation](https://velero.io/docs/main/) to add schedules, volume snapshots, and cluster migration flows. diff --git a/content/fr/developer/integration/big-data/hudi.md b/content/fr/developer/integration/big-data/hudi.md new file mode 100644 index 00000000..42026454 --- /dev/null +++ b/content/fr/developer/integration/big-data/hudi.md @@ -0,0 +1,109 @@ +--- +title: "Apache Hudi" +description: "Write Apache Hudi tables to RustFS through Spark and the s3a connector." +--- + +This guide connects [Apache Hudi](https://github.com/apache/hudi) — the transactional data lake platform — to **RustFS** as its copy-on-write table store. You will run Spark with the Hudi bundle, write a table to the `s3a://` location inside a RustFS bucket, read it back, and verify the table files. The workflow was verified with `apache/spark:3.5.6`, `hudi-spark3.5-bundle_2.12:0.15.0`, `hadoop-aws:3.3.4`, and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Job["Spark + Hudi"] -->|"commit + parquet"| RustFS["RustFS :9000"] +``` + +Hudi stores the `.hoodie/` timeline, commit files, and Parquet data blocks under the table path in the bucket. Reads resolve the latest table snapshot from the timeline, so every write and query goes through the S3 API. + +## 1. Run the Spark write + +Hudi requires the Kryo serializer and the s3a credentials as Hadoop properties. The following `spark-shell` session creates a non-partitioned table in `my-bucket` — replace all connection placeholders: + +```scala +import org.apache.spark.sql.SaveMode + +val df = Seq((1, "a"), (2, "b"), (3, "c")).toDF("id", "name") +df.write.format("hudi") + .option("hoodie.table.name", "events") + .option("hoodie.datasource.write.recordkey.field", "id") + .option("hoodie.datasource.write.precombine.field", "name") + .option("hoodie.datasource.write.partitionpath.field", "") + .mode(SaveMode.Overwrite) + .save("s3a://my-bucket/hudi-demo/events") + +val back = spark.read.format("hudi").load("s3a://my-bucket/hudi-demo/events") +back.select("id", "name").show() +``` + +```bash +docker run --rm --network oo-rustfs_default \ + -v "$PWD/hudi_test.scala":/tmp/hudi_test.scala \ + apache/spark:3.5.6 /opt/spark/bin/spark-shell \ + --packages org.apache.hudi:hudi-spark3.5-bundle_2.12:0.15.0,org.apache.hadoop:hadoop-aws:3.3.4 \ + --conf spark.serializer=org.apache.spark.serializer.KryoSerializer \ + --conf spark.hadoop.fs.s3a.endpoint=http://rustfs:9000 \ + --conf spark.hadoop.fs.s3a.path.style.access=true \ + --conf spark.hadoop.fs.s3a.access.key= \ + --conf spark.hadoop.fs.s3a.secret.key= \ + --conf spark.hadoop.fs.s3a.region=us-east-1 \ + --conf spark.jars.ivy=/tmp/.ivy2 \ + --conf spark.sql.shuffle.partitions=2 \ + -i /tmp/hudi_test.scala +``` + +The `--packages` flags download the Hudi bundle and the s3a connector on first run. The `spark.jars.ivy` setting avoids Ivy cache permission errors in the container. The read-back `show()` prints the three rows: + +```text ++---+----+ +| 2| b| +| 3| c| +| 1| a| ++---+----+ +``` + +## 2. Verify objects in RustFS + +List the table prefix: + +```bash +rc ls rustfs/my-bucket/hudi-demo/ -r +``` + +The Hudi table layout appears under the table path — the `.hoodie/` timeline with the commit file, plus the Parquet data block: + +```text +hudi-demo/events/.hoodie/20260922015803291.commit +hudi-demo/events/.hoodie/20260922015803291.commit.requested +hudi-demo/events/// +``` + +![Hudi table files stored in the RustFS Console](./images/rustfs-hudi-table.png) + +## 3. Stop or reset + +Spark runs as a one-shot client and holds no state. To delete the demo table: + +```bash +rc rm rustfs/my-bucket/hudi-demo/ --recursive --force +``` + +## Troubleshooting + +### `hoodie only support org.apache.spark.serializer.KryoSerializer as spark.serializer` + +Hudi rejects the default Spark serializer. Pass `--conf spark.serializer=org.apache.spark.serializer.KryoSerializer` as shown above. + +### `Partition-path field has to be non-empty` or keygenerator class errors + +For a non-partitioned table, keep the default key generator (do not set `hoodie.datasource.write.keygenerator.class`) and set `hoodie.datasource.write.partitionpath.field` to the empty string. + +### `NoSuchMethodError` or `ClassNotFoundException` in the Hudi write + +The Spark and Hudi bundle versions must match: `hudi-spark3.5-bundle_2.12` goes with Spark 3.5.x, and the bundled Hadoop client must be compatible with `hadoop-aws:3.3.4`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Hudi operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Hudi Spark guide](https://hudi.apache.org/docs/quick-start-guide/) to add upserts, compaction, and query integrations on the same table. diff --git a/content/fr/developer/integration/big-data/images/rustfs-hudi-table.png b/content/fr/developer/integration/big-data/images/rustfs-hudi-table.png new file mode 100644 index 00000000..b91c9157 Binary files /dev/null and b/content/fr/developer/integration/big-data/images/rustfs-hudi-table.png differ diff --git a/content/fr/developer/integration/big-data/images/rustfs-lakefs-repo.png b/content/fr/developer/integration/big-data/images/rustfs-lakefs-repo.png new file mode 100644 index 00000000..0236a479 Binary files /dev/null and b/content/fr/developer/integration/big-data/images/rustfs-lakefs-repo.png differ diff --git a/content/fr/developer/integration/big-data/index.md b/content/fr/developer/integration/big-data/index.md index 41f295dd..add692d8 100644 --- a/content/fr/developer/integration/big-data/index.md +++ b/content/fr/developer/integration/big-data/index.md @@ -8,6 +8,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo ## Systems - [ClickHouse](./clickhouse.md) +- [Hudi](./hudi.md) - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) - [Milvus](./milvus.md) @@ -15,6 +16,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [OpenDAL](./opendal.md) - [DuckDB](./duckdb.md) - [Doris](./doris.md) +- [lakeFS](./lakefs.md) - [InfluxDB](./influxdb.md) - [Spark](./spark.md) - [Flink](./flink.md) diff --git a/content/fr/developer/integration/big-data/lakefs.md b/content/fr/developer/integration/big-data/lakefs.md new file mode 100644 index 00000000..ba63ef22 --- /dev/null +++ b/content/fr/developer/integration/big-data/lakefs.md @@ -0,0 +1,160 @@ +--- +title: "lakeFS" +description: "Run lakeFS with RustFS as its S3 blockstore for versioned data lakes." +--- + +This guide connects [lakeFS](https://github.com/treeverse/lakeFS) — the Git-like data lake versioning layer — to **RustFS** as its S3 blockstore. You will start lakeFS, create a repository whose storage namespace points at a RustFS bucket, commit an object, and verify that the lakeFS metadata and data files live in RustFS. The workflow was verified with `treeverse/lakefs:1.58.0` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["lakectl / API"] -->|HTTP| LakeFS["lakeFS :8000"] + LakeFS -->|"metadata + data"| RustFS["RustFS :9000"] +``` + +lakeFS stores repository metadata and committed data files in the blockstore under the storage namespace. The bucket holds one `repo/` prefix containing `_lakefs/` metadata and `data/` objects; lakeFS reads and writes them through the S3 API. + +## 1. Create the project files + +Create the bucket, then the lakeFS configuration: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/lakefs-data +``` + +```yaml title="config.yaml" +database: + type: local + local: + path: /lakefs/data + +blockstore: + type: s3 + s3: + endpoint: http://rustfs:9000 + region: us-east-1 + force_path_style: true + +auth: + encrypt: + secret_key: change-me-to-a-random-string + +gateways: + s3: + domain_name: rustfs-gateway.local + +logging: + format: text + level: info +``` + +`force_path_style: true` is required — without it lakeFS builds `lakefs-data.` as a hostname and every request fails with a DNS error. The credentials are supplied through the standard `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` environment variables. + +Create the environment file for the service credentials and the initial admin user, replacing the placeholders: + +```ini title=".env" +AWS_ACCESS_KEY_ID= +AWS_SECRET_ACCESS_KEY= +LAKEFS_INSTALLATION_USER_NAME=admin +LAKEFS_INSTALLATION_ACCESS_KEY_ID= +LAKEFS_INSTALLATION_SECRET_ACCESS_KEY= +LAKEFS_STATS_ENABLED=false +``` + +## 2. Start lakeFS + +```bash +docker run -d --name lakefs --network oo-rustfs_default \ + -p 8000:8000 \ + -v "$PWD/config.yaml":/etc/lakefs/config.yaml:ro \ + --env-file .env \ + treeverse/lakefs:1.58.0 run +``` + +Wait for `http://localhost:8000/api/healthcheck` to return `200`, then create a repository with a storage namespace inside the bucket: + +```bash +curl -s -u : \ + -X POST http://localhost:8000/api/v1/repositories \ + -H "Content-Type: application/json" \ + -d '{"name": "rustfs-demo", "storage_namespace": "s3://lakefs-data/repo"}' +``` + +## 3. Commit an object + +Upload a file to the `main` branch and commit it: + +```bash +echo "hello from lakefs on rustfs" > hello.txt + +curl -s -u : \ + -X POST "http://localhost:8000/api/v1/repositories/rustfs-demo/branches/main/objects?path=hello.txt" \ + --data-binary @hello.txt + +curl -s -u : \ + -X POST "http://localhost:8000/api/v1/repositories/rustfs-demo/branches/main/commits" \ + -H "Content-Type: application/json" \ + -d '{"message": "add hello"}' +``` + +Read the object back through the branch ref — the response body is the committed content: + +```bash +curl -s -u : \ + "http://localhost:8000/api/v1/repositories/rustfs-demo/refs/main/objects?path=hello.txt" +``` + +## 4. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/lakefs-data/ -r +``` + +The output shows the lakeFS metadata objects and the committed data file under the repository prefix: + +```text +repo/_lakefs/19b2b26e37cb20fc6763c527f88eb5151891b04a2c8c9ddd32870c5c3f353281 +repo/data/fueia10jdra000e1c480/daots7ojdra000e1c490 +``` + +![lakeFS objects stored in the RustFS Console](./images/rustfs-lakefs-repo.png) + +## 5. Stop or reset the deployment + +Stop lakeFS while keeping the data: + +```bash +docker rm -f lakefs +``` + +The repository metadata and data stay in the `lakefs-data` bucket, so restarting lakeFS with the same configuration brings the repository back. To delete everything, remove the bucket: + +```bash +rc rb rustfs/lakefs-data --force +``` + +## Troubleshooting + +### `failed to create repository: failed to access storage` with a DNS error + +lakeFS is using virtual-hosted addressing. Set `force_path_style: true` under `blockstore.s3` — in the 1.x configuration schema the older `path_style` key is rejected at startup. + +### `missing required keys: [auth.encrypt.secret_key]` + +lakeFS 1.x requires an encryption key for the local database. Add the `auth.encrypt.secret_key` block as shown in the configuration above. + +### `mkdir /lakefs: permission denied` + +The container runs as a non-root user and cannot create the database directory. Run the container with `-u 0` for local testing, or mount a writable directory at the configured path. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional lakeFS operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [lakeFS S3 blockstore documentation](https://docs.lakefs.io/howto/using-s3.html) to configure the S3 gateway for tools that speak the S3 protocol. diff --git a/content/fr/developer/integration/big-data/meta.json b/content/fr/developer/integration/big-data/meta.json index ee7d69b5..249e219d 100644 --- a/content/fr/developer/integration/big-data/meta.json +++ b/content/fr/developer/integration/big-data/meta.json @@ -2,16 +2,18 @@ "title": "Data Analytics", "pages": [ "clickhouse", + "duckdb", + "doris", + "flink", + "hudi", "iceberg", - "pyiceberg", + "influxdb", + "lakefs", "milvus", "mlflow", "opendal", - "duckdb", - "doris", - "influxdb", + "pyiceberg", "spark", - "flink", "trino", "zeppelin" ] diff --git a/content/fr/developer/integration/index.md b/content/fr/developer/integration/index.md index c7404168..19dfbdf8 100644 --- a/content/fr/developer/integration/index.md +++ b/content/fr/developer/integration/index.md @@ -8,8 +8,9 @@ Utilisez cette section pour connecter **RustFS** à des plateformes d'infrastruc ## Integration categories - [Reverse Proxy](./reverse-proxy/index.md) couvre Nginx, Traefik, Caddy et HAProxy. -- [Backup](./backup/index.md) couvre Restic et Longhorn. -- [Analyse de données](./big-data/index.md) couvre les systèmes d'analyse incluant ClickHouse, Doris, Iceberg, Milvus, OpenDAL et Zeppelin. +- [Backup](./backup/index.md) couvre Kopia, Longhorn, Restic et Velero. +- [IA](./ai/index.md) couvre les plateformes d'IA incluant Ray. +- [Analyse de données](./big-data/index.md) couvre les systèmes d'analyse incluant ClickHouse, Doris, Hudi, Iceberg, lakeFS, Milvus, OpenDAL et Zeppelin. - [Observabilité](./observability/index.md) couvre les systèmes de télémétrie incluant Fluentd, OpenObserve, OpenTelemetry, Thanos et Tempo. - [Autres](./others/index.md) couvre le SDK communautaire capo pour Python. - [Registre](./registry/index.md) couvre Harbor. diff --git a/content/fr/developer/integration/meta.json b/content/fr/developer/integration/meta.json index b4b29480..a3c989a5 100644 --- a/content/fr/developer/integration/meta.json +++ b/content/fr/developer/integration/meta.json @@ -4,6 +4,7 @@ "reverse-proxy", "backup", "big-data", + "ai", "observability", "others", "registry", diff --git a/content/ja/developer/integration/ai/images/rustfs-ray-data.png b/content/ja/developer/integration/ai/images/rustfs-ray-data.png new file mode 100644 index 00000000..394f4499 Binary files /dev/null and b/content/ja/developer/integration/ai/images/rustfs-ray-data.png differ diff --git a/content/ja/developer/integration/ai/index.md b/content/ja/developer/integration/ai/index.md new file mode 100644 index 00000000..fb36b0fc --- /dev/null +++ b/content/ja/developer/integration/ai/index.md @@ -0,0 +1,12 @@ +--- +title: "AI" +description: "S3 互換オブジェクトストレージインターフェースを経由して、AI プラットフォームを RustFS に接続します。" +--- + +S3 互換エンドポイントをサポートする AI・機械学習プラットフォームのオブジェクトストレージ層として **RustFS** を使用します。 + +## プラットフォーム + +- [Ray](./ray.md) + +トレーニングデータセットとチェックポイントは専用バケットに保存し、必要なバケット操作のみに権限が絞られた認証情報を使用してください。 diff --git a/content/ja/developer/integration/ai/meta.json b/content/ja/developer/integration/ai/meta.json new file mode 100644 index 00000000..573fc520 --- /dev/null +++ b/content/ja/developer/integration/ai/meta.json @@ -0,0 +1,6 @@ +{ + "title": "AI", + "pages": [ + "ray" + ] +} diff --git a/content/ja/developer/integration/ai/ray.md b/content/ja/developer/integration/ai/ray.md new file mode 100644 index 00000000..e9f71b94 --- /dev/null +++ b/content/ja/developer/integration/ai/ray.md @@ -0,0 +1,103 @@ +--- +title: "Ray" +description: "Use Ray Data with RustFS as S3-compatible storage for dataset writes and reads." +--- + +This guide connects [Ray](https://github.com/ray-project/ray) — the distributed AI and Python compute framework — to **RustFS** through Ray Data's S3 filesystem support. You will run a Ray job inside the official image, write a dataset as Parquet to a RustFS bucket, read it back, and verify the objects. The workflow was verified with `rayproject/ray:2.44.0-py311` (Ray 2.44, pyarrow filesystem) and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Job["Ray job"] -->|"ray.data"| DS["Dataset"] + DS -->|"Parquet files"| RustFS["RustFS :9000"] +``` + +Ray Data reads and writes datasets through pyarrow's `S3FileSystem`. Passing an `S3FileSystem` configured for RustFS redirects every dataset operation — Parquet, CSV, JSON — to the bucket. + +## 1. Create the job file + +Create the script, replacing all connection placeholders: + +```python title="ray_s3.py" +import ray +ray.init(ignore_reinit_error=True) + +import pandas as pd +from pyarrow.fs import S3FileSystem + +fs = S3FileSystem( + endpoint_override="http://:9000", + access_key="", + secret_key="", + region="us-east-1", +) + +df = pd.DataFrame({"id": range(5), "value": [x * 1.5 for x in range(5)]}) +ds = ray.data.from_pandas(df) +ds.write_parquet("my-bucket/ray-demo/events/", filesystem=fs) + +back = ray.data.read_parquet("my-bucket/ray-demo/events/", filesystem=fs).take_all() +print("rows:", len(back)) +print("sample:", back[0]) +ray.shutdown() +``` + +`endpoint_override` takes the full endpoint URL including the scheme. pyarrow's `S3FileSystem` uses path-style requests for custom endpoints, so no extra flag is needed. The same filesystem object works for `write_csv`, `read_json`, and the other Ray Data methods. + +## 2. Run the job + +Run the script in the Ray image on the same Docker network as RustFS: + +```bash +docker run --rm --network oo-rustfs_default \ + -v "$PWD/ray_s3.py":/tmp/ray_s3.py \ + rayproject/ray:2.44.0-py311 python /tmp/ray_s3.py +``` + +```text +rows: 5 +sample: {'id': 0, 'value': 0.0} +``` + +## 3. Verify objects in RustFS + +List the dataset prefix: + +```bash +rc ls rustfs/my-bucket/ray-demo/ -r +``` + +Ray Data wrote the dataset as a Parquet block: + +```text +ray-demo/events/0_000000_000000.parquet +``` + +![Ray dataset files stored in the RustFS Console](./images/rustfs-ray-data.png) + +## 4. Stop or reset + +Ray Data holds no state of its own. To delete the demo dataset: + +```bash +rc rm rustfs/my-bucket/ray-demo/ --recursive --force +``` + +## Troubleshooting + +### `Unable to connect to endpoint` or timeouts + +Confirm `endpoint_override` includes the scheme and is reachable from the Ray container. Inside a Compose network the hostname is `rustfs`; from the host use `http://localhost:9000`. + +### `Access Denied` on write + +Confirm the access key and secret key are passed to `S3FileSystem` itself — Ray does not read the container's AWS environment variables through this code path. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Ray operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Ray Data documentation](https://docs.ray.io/en/latest/data/data.html) to chain transformations, training ingestion, and checkpointing on the same bucket. diff --git a/content/ja/developer/integration/backup/images/rustfs-kopia-repo.png b/content/ja/developer/integration/backup/images/rustfs-kopia-repo.png new file mode 100644 index 00000000..4241891c Binary files /dev/null and b/content/ja/developer/integration/backup/images/rustfs-kopia-repo.png differ diff --git a/content/ja/developer/integration/backup/images/rustfs-velero-backups.png b/content/ja/developer/integration/backup/images/rustfs-velero-backups.png new file mode 100644 index 00000000..c3e9b169 Binary files /dev/null and b/content/ja/developer/integration/backup/images/rustfs-velero-backups.png differ diff --git a/content/ja/developer/integration/backup/index.md b/content/ja/developer/integration/backup/index.md index 63f7958a..a63d6b28 100644 --- a/content/ja/developer/integration/backup/index.md +++ b/content/ja/developer/integration/backup/index.md @@ -8,6 +8,8 @@ description: "S3 互換のオブジェクトストレージ経由でバックア ## システム - [Restic](./restic.md) +- [Velero](./velero.md) +- [Kopia](./kopia.md) - [Longhorn](./longhorn.md) バックアップジョブは専用のバケットとプレフィックスにまとめ、必要なバケット操作だけに絞った認証情報を使用します。 \ No newline at end of file diff --git a/content/ja/developer/integration/backup/kopia.md b/content/ja/developer/integration/backup/kopia.md new file mode 100644 index 00000000..fed21e5a --- /dev/null +++ b/content/ja/developer/integration/backup/kopia.md @@ -0,0 +1,125 @@ +--- +title: "Kopia" +description: "Back up files to RustFS with Kopia's S3 repository backend." +--- + +This guide connects [Kopia](https://github.com/kopia/kopia) — the open-source backup and restore tool — to **RustFS** as an S3 repository. You will create a repository in a RustFS bucket, take a snapshot of a directory, restore it into an empty directory, and compare checksums. The workflow was verified with `kopia/kopia:0.18.1` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Source["Source files"] -->|snapshot| Kopia["Kopia"] + Kopia -->|"encrypted blocks"| RustFS["RustFS :9000"] + Kopia -->|restore| Restore["Restored files"] +``` + +Kopia stores the repository format files and deduplicated, encrypted content blocks in the bucket. Restores read the blocks back and reassemble the original files, so the checksum of every restored file must match the source. + +## 1. Create the repository + +Create the bucket first, then initialize the Kopia repository inside it. The endpoint is a bare `host:port` (no scheme); `--disable-tls` switches the client to plain HTTP: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/kopia-backups + +docker run --rm --network oo-rustfs_default kopia/kopia:0.18.1 repository create s3 \ + --bucket kopia-backups \ + --access-key \ + --secret-access-key \ + --endpoint :9000 \ + --region us-east-1 \ + --disable-tls \ + --password \ + --override-username demo --override-hostname workstation +``` + +Kopia validates the provider by reading and writing through the S3 API before it reports success. + +## 2. Connect, snapshot, and restore + +Run the following commands from the directory holding your `repository.config` (created by the previous step). Kopia reads the connection settings from that file, so the S3 flags are only needed once: + +```bash +export KOPIA_PASSWORD= +export KOPIA_CONFIG_PATH=/config/repository.config + +alias kopia='docker run --rm --network oo-rustfs_default \ + -e KOPIA_PASSWORD -e KOPIA_CONFIG_PATH \ + -v "$PWD/config:/config" \ + -v "$PWD/source:/source:ro" \ + -v "$PWD/restore:/restore" kopia/kopia:0.18.1' + +kopia repository connect s3 \ + --bucket kopia-backups \ + --access-key \ + --secret-access-key \ + --endpoint :9000 \ + --region us-east-1 --disable-tls \ + --override-username demo --override-hostname workstation + +kopia snapshot create /source +kopia snapshot list +``` + +Restore the snapshot into an empty directory and compare checksums with the source. The snapshot ID is the `ka...` identifier printed by `snapshot list`: + +```bash +kopia restore /restore + +sha256sum source/blob.bin restore/blob.bin +``` + +```text +bec4530e2798465b... source/blob.bin +bec4530e2798465b... restore/blob.bin +``` + +## 3. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/kopia-backups/ -r +``` + +The output shows the repository format files plus the packed content blocks written by the snapshot: + +```text +kopia.blobcfg +kopia.repository +p0000.../... +``` + +![Kopia repository blocks stored in the RustFS Console](./images/rustfs-kopia-repo.png) + +## 4. Stop or reset + +Kopia is a client-side tool and holds no running state. To delete the repository and all snapshots, remove the bucket: + +```bash +rc rb rustfs/kopia-backups --force +``` + +## Troubleshooting + +### `Endpoint url cannot have fully qualified paths` + +The endpoint must be a bare `host:port` value without a scheme or path — Kopia builds the object URLs itself. + +### `server gave HTTP response to HTTPS client` + +Without `--disable-tls`, Kopia speaks HTTPS. RustFS without TLS needs the `--disable-tls` flag on both `repository create s3` and `repository connect s3`. + +### `can't connect to storage` with a DNS error + +The S3 client is using virtual-hosted addressing. Keep the endpoint as a bare `host:port` value; Kopia uses path-style requests for endpoints given in that form. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Kopia operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Kopia repository documentation](https://kopia.io/docs/repositories/) to add policies, retention, and scheduled snapshots. diff --git a/content/ja/developer/integration/backup/meta.json b/content/ja/developer/integration/backup/meta.json index 6cd71cbf..e7421b9c 100644 --- a/content/ja/developer/integration/backup/meta.json +++ b/content/ja/developer/integration/backup/meta.json @@ -1,7 +1,9 @@ { "title": "バックアップ", "pages": [ + "kopia", + "longhorn", "restic", - "longhorn" + "velero" ] } diff --git a/content/ja/developer/integration/backup/velero.md b/content/ja/developer/integration/backup/velero.md new file mode 100644 index 00000000..0140296b --- /dev/null +++ b/content/ja/developer/integration/backup/velero.md @@ -0,0 +1,153 @@ +--- +title: "Velero" +description: "Back up Kubernetes cluster resources to RustFS with Velero's AWS object store provider." +--- + +This guide connects [Velero](https://github.com/vmware-tanzu/velero) — the Kubernetes backup and restore tool — to **RustFS** as its object storage backend. You will install Velero into a Kubernetes cluster with a RustFS backup location, back up cluster resources, delete them, restore from the bucket, and verify the restored objects. The workflow was verified with Velero CLI 1.16.2, `velero-plugin-for-aws:v1.12.2` on k3s (Kubernetes 1.30), and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need a Kubernetes cluster with `kubectl` access and the Velero CLI. This guide is intended for integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + K8s["Kubernetes cluster"] -->|"resources"| V["Velero"] + V -->|"backups + logs"| RustFS["RustFS :9000"] +``` + +Velero serializes Kubernetes resources and (optionally) pod volume data into gzip archives under `backups//` in the bucket. Restores download those archives and recreate the resources in the cluster. + +## 1. Create the backup bucket + +Create a dedicated bucket — Velero does not create buckets: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/velero-backups +``` + +## 2. Install Velero + +Write the credentials to a file and install the Velero server. The `s3Url` must be an endpoint reachable from the cluster pods themselves — use the node IP or an internal address, not a port forward from your workstation: + +```bash +cat > velero-creds < +aws_secret_access_key= +EOF + +kubectl create namespace velero +kubectl create secret generic cloud-credentials \ + --namespace velero --from-file=cloud=velero-creds + +velero install \ + --provider aws \ + --plugins velero/velero-plugin-for-aws:v1.12.2 \ + --bucket velero-backups \ + --backup-location-config region=us-east-1,s3ForcePathStyle="true",s3Url=http://:9000 \ + --secret-file velero-creds \ + --use-volume-snapshots=false +``` + +Wait for the deployment to become ready and the backup location to turn `Available`: + +```bash +kubectl -n velero get pods +velero backup-location get +``` + +```text +NAME PROVIDER BUCKET/PREFIX PHASE LAST VALIDATED ACCESS MODE DEFAULT +default aws velero-backups Available 2026-09-22 ... ReadWrite true +``` + +## 3. Back up cluster resources + +Create two demo resources and back up the whole default namespace scope: + +```bash +kubectl create configmap demo-cm --from-literal=key=rustfs-velero-demo +kubectl create deployment nginx --image=nginx:1.27 + +velero backup create demo-backup --wait +velero backup get +``` + +The backup uploads the resource archives and logs to RustFS. + +## 4. Verify the backup in RustFS + +List the bucket: + +```bash +rc ls rustfs/velero-backups/ -r +``` + +The backup archive set is stored under the `backups/` prefix: + +```text +backups/demo-backup/demo-backup-resources.json.gz +backups/demo-backup/demo-backup-logs.gz +backups/demo-backup/demo-backup-itemoperations.json.gz +``` + +![Velero backups stored in the RustFS Console](./images/rustfs-velero-backups.png) + +## 5. Restore and verify + +Delete the demo resources and restore them from the backup. The restore reads the archives from RustFS and recreates the resources: + +```bash +kubectl delete configmap demo-cm +kubectl delete deployment nginx + +velero restore create --from-backup demo-backup --wait +``` + +Confirm the resources are back with their original content: + +```bash +kubectl get configmap demo-cm -o jsonpath="{.data.key}" +kubectl get deployment nginx +``` + +```text +rustfs-velero-demo +NAME READY UP-TO-DATE AVAILABLE AGE +nginx 1/1 1 1 10s +``` + +## 6. Stop or reset + +The backups stay in the `velero-backups` bucket and restore into any cluster that runs Velero against the same bucket. To remove the demo backup from RustFS: + +```bash +velero backup delete demo-backup --confirm +rc rm rustfs/velero-backups/backups/ --recursive --force +``` + +## Troubleshooting + +### The backup location stays `Unavailable` + +The Velero pod cannot reach the `s3Url`. Endpoints on a Docker custom bridge network are not routable from Kubernetes pods — use the Docker host bridge IP (for example `http://172.17.0.1:9000` on a single-host cluster) or a node address. Patch the location and let Velero revalidate: + +```bash +kubectl -n velero patch backupstoragelocation default --type merge -p \ + '{"spec":{"config":{"s3Url":"http://172.17.0.1:9000"}}}' +``` + +### The backup ends `PartiallyFailed` with a daemonset error + +The error `daemonset pod not found in running state` means the node-agent pod (used for pod volume backups) was not running. It does not affect the cluster-resource archives. Wait for the agent or keep `--default-volumes-to-fs-backup` only when the agent is running. + +### `FailedValidation` right after install + +The first backup attempt runs against a location that has not validated yet. Wait for `velero backup-location get` to report `Available` before creating backups. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Velero operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Velero documentation](https://velero.io/docs/main/) to add schedules, volume snapshots, and cluster migration flows. diff --git a/content/ja/developer/integration/big-data/hudi.md b/content/ja/developer/integration/big-data/hudi.md new file mode 100644 index 00000000..42026454 --- /dev/null +++ b/content/ja/developer/integration/big-data/hudi.md @@ -0,0 +1,109 @@ +--- +title: "Apache Hudi" +description: "Write Apache Hudi tables to RustFS through Spark and the s3a connector." +--- + +This guide connects [Apache Hudi](https://github.com/apache/hudi) — the transactional data lake platform — to **RustFS** as its copy-on-write table store. You will run Spark with the Hudi bundle, write a table to the `s3a://` location inside a RustFS bucket, read it back, and verify the table files. The workflow was verified with `apache/spark:3.5.6`, `hudi-spark3.5-bundle_2.12:0.15.0`, `hadoop-aws:3.3.4`, and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Job["Spark + Hudi"] -->|"commit + parquet"| RustFS["RustFS :9000"] +``` + +Hudi stores the `.hoodie/` timeline, commit files, and Parquet data blocks under the table path in the bucket. Reads resolve the latest table snapshot from the timeline, so every write and query goes through the S3 API. + +## 1. Run the Spark write + +Hudi requires the Kryo serializer and the s3a credentials as Hadoop properties. The following `spark-shell` session creates a non-partitioned table in `my-bucket` — replace all connection placeholders: + +```scala +import org.apache.spark.sql.SaveMode + +val df = Seq((1, "a"), (2, "b"), (3, "c")).toDF("id", "name") +df.write.format("hudi") + .option("hoodie.table.name", "events") + .option("hoodie.datasource.write.recordkey.field", "id") + .option("hoodie.datasource.write.precombine.field", "name") + .option("hoodie.datasource.write.partitionpath.field", "") + .mode(SaveMode.Overwrite) + .save("s3a://my-bucket/hudi-demo/events") + +val back = spark.read.format("hudi").load("s3a://my-bucket/hudi-demo/events") +back.select("id", "name").show() +``` + +```bash +docker run --rm --network oo-rustfs_default \ + -v "$PWD/hudi_test.scala":/tmp/hudi_test.scala \ + apache/spark:3.5.6 /opt/spark/bin/spark-shell \ + --packages org.apache.hudi:hudi-spark3.5-bundle_2.12:0.15.0,org.apache.hadoop:hadoop-aws:3.3.4 \ + --conf spark.serializer=org.apache.spark.serializer.KryoSerializer \ + --conf spark.hadoop.fs.s3a.endpoint=http://rustfs:9000 \ + --conf spark.hadoop.fs.s3a.path.style.access=true \ + --conf spark.hadoop.fs.s3a.access.key= \ + --conf spark.hadoop.fs.s3a.secret.key= \ + --conf spark.hadoop.fs.s3a.region=us-east-1 \ + --conf spark.jars.ivy=/tmp/.ivy2 \ + --conf spark.sql.shuffle.partitions=2 \ + -i /tmp/hudi_test.scala +``` + +The `--packages` flags download the Hudi bundle and the s3a connector on first run. The `spark.jars.ivy` setting avoids Ivy cache permission errors in the container. The read-back `show()` prints the three rows: + +```text ++---+----+ +| 2| b| +| 3| c| +| 1| a| ++---+----+ +``` + +## 2. Verify objects in RustFS + +List the table prefix: + +```bash +rc ls rustfs/my-bucket/hudi-demo/ -r +``` + +The Hudi table layout appears under the table path — the `.hoodie/` timeline with the commit file, plus the Parquet data block: + +```text +hudi-demo/events/.hoodie/20260922015803291.commit +hudi-demo/events/.hoodie/20260922015803291.commit.requested +hudi-demo/events/// +``` + +![Hudi table files stored in the RustFS Console](./images/rustfs-hudi-table.png) + +## 3. Stop or reset + +Spark runs as a one-shot client and holds no state. To delete the demo table: + +```bash +rc rm rustfs/my-bucket/hudi-demo/ --recursive --force +``` + +## Troubleshooting + +### `hoodie only support org.apache.spark.serializer.KryoSerializer as spark.serializer` + +Hudi rejects the default Spark serializer. Pass `--conf spark.serializer=org.apache.spark.serializer.KryoSerializer` as shown above. + +### `Partition-path field has to be non-empty` or keygenerator class errors + +For a non-partitioned table, keep the default key generator (do not set `hoodie.datasource.write.keygenerator.class`) and set `hoodie.datasource.write.partitionpath.field` to the empty string. + +### `NoSuchMethodError` or `ClassNotFoundException` in the Hudi write + +The Spark and Hudi bundle versions must match: `hudi-spark3.5-bundle_2.12` goes with Spark 3.5.x, and the bundled Hadoop client must be compatible with `hadoop-aws:3.3.4`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Hudi operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Hudi Spark guide](https://hudi.apache.org/docs/quick-start-guide/) to add upserts, compaction, and query integrations on the same table. diff --git a/content/ja/developer/integration/big-data/images/rustfs-hudi-table.png b/content/ja/developer/integration/big-data/images/rustfs-hudi-table.png new file mode 100644 index 00000000..b91c9157 Binary files /dev/null and b/content/ja/developer/integration/big-data/images/rustfs-hudi-table.png differ diff --git a/content/ja/developer/integration/big-data/images/rustfs-lakefs-repo.png b/content/ja/developer/integration/big-data/images/rustfs-lakefs-repo.png new file mode 100644 index 00000000..0236a479 Binary files /dev/null and b/content/ja/developer/integration/big-data/images/rustfs-lakefs-repo.png differ diff --git a/content/ja/developer/integration/big-data/index.md b/content/ja/developer/integration/big-data/index.md index cca193e0..cd8c7ce7 100644 --- a/content/ja/developer/integration/big-data/index.md +++ b/content/ja/developer/integration/big-data/index.md @@ -8,6 +8,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo ## Systems - [ClickHouse](./clickhouse.md) +- [Hudi](./hudi.md) - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) - [Milvus](./milvus.md) @@ -15,6 +16,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [OpenDAL](./opendal.md) - [DuckDB](./duckdb.md) - [Doris](./doris.md) +- [lakeFS](./lakefs.md) - [InfluxDB](./influxdb.md) - [Spark](./spark.md) - [Flink](./flink.md) diff --git a/content/ja/developer/integration/big-data/lakefs.md b/content/ja/developer/integration/big-data/lakefs.md new file mode 100644 index 00000000..ba63ef22 --- /dev/null +++ b/content/ja/developer/integration/big-data/lakefs.md @@ -0,0 +1,160 @@ +--- +title: "lakeFS" +description: "Run lakeFS with RustFS as its S3 blockstore for versioned data lakes." +--- + +This guide connects [lakeFS](https://github.com/treeverse/lakeFS) — the Git-like data lake versioning layer — to **RustFS** as its S3 blockstore. You will start lakeFS, create a repository whose storage namespace points at a RustFS bucket, commit an object, and verify that the lakeFS metadata and data files live in RustFS. The workflow was verified with `treeverse/lakefs:1.58.0` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["lakectl / API"] -->|HTTP| LakeFS["lakeFS :8000"] + LakeFS -->|"metadata + data"| RustFS["RustFS :9000"] +``` + +lakeFS stores repository metadata and committed data files in the blockstore under the storage namespace. The bucket holds one `repo/` prefix containing `_lakefs/` metadata and `data/` objects; lakeFS reads and writes them through the S3 API. + +## 1. Create the project files + +Create the bucket, then the lakeFS configuration: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/lakefs-data +``` + +```yaml title="config.yaml" +database: + type: local + local: + path: /lakefs/data + +blockstore: + type: s3 + s3: + endpoint: http://rustfs:9000 + region: us-east-1 + force_path_style: true + +auth: + encrypt: + secret_key: change-me-to-a-random-string + +gateways: + s3: + domain_name: rustfs-gateway.local + +logging: + format: text + level: info +``` + +`force_path_style: true` is required — without it lakeFS builds `lakefs-data.` as a hostname and every request fails with a DNS error. The credentials are supplied through the standard `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` environment variables. + +Create the environment file for the service credentials and the initial admin user, replacing the placeholders: + +```ini title=".env" +AWS_ACCESS_KEY_ID= +AWS_SECRET_ACCESS_KEY= +LAKEFS_INSTALLATION_USER_NAME=admin +LAKEFS_INSTALLATION_ACCESS_KEY_ID= +LAKEFS_INSTALLATION_SECRET_ACCESS_KEY= +LAKEFS_STATS_ENABLED=false +``` + +## 2. Start lakeFS + +```bash +docker run -d --name lakefs --network oo-rustfs_default \ + -p 8000:8000 \ + -v "$PWD/config.yaml":/etc/lakefs/config.yaml:ro \ + --env-file .env \ + treeverse/lakefs:1.58.0 run +``` + +Wait for `http://localhost:8000/api/healthcheck` to return `200`, then create a repository with a storage namespace inside the bucket: + +```bash +curl -s -u : \ + -X POST http://localhost:8000/api/v1/repositories \ + -H "Content-Type: application/json" \ + -d '{"name": "rustfs-demo", "storage_namespace": "s3://lakefs-data/repo"}' +``` + +## 3. Commit an object + +Upload a file to the `main` branch and commit it: + +```bash +echo "hello from lakefs on rustfs" > hello.txt + +curl -s -u : \ + -X POST "http://localhost:8000/api/v1/repositories/rustfs-demo/branches/main/objects?path=hello.txt" \ + --data-binary @hello.txt + +curl -s -u : \ + -X POST "http://localhost:8000/api/v1/repositories/rustfs-demo/branches/main/commits" \ + -H "Content-Type: application/json" \ + -d '{"message": "add hello"}' +``` + +Read the object back through the branch ref — the response body is the committed content: + +```bash +curl -s -u : \ + "http://localhost:8000/api/v1/repositories/rustfs-demo/refs/main/objects?path=hello.txt" +``` + +## 4. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/lakefs-data/ -r +``` + +The output shows the lakeFS metadata objects and the committed data file under the repository prefix: + +```text +repo/_lakefs/19b2b26e37cb20fc6763c527f88eb5151891b04a2c8c9ddd32870c5c3f353281 +repo/data/fueia10jdra000e1c480/daots7ojdra000e1c490 +``` + +![lakeFS objects stored in the RustFS Console](./images/rustfs-lakefs-repo.png) + +## 5. Stop or reset the deployment + +Stop lakeFS while keeping the data: + +```bash +docker rm -f lakefs +``` + +The repository metadata and data stay in the `lakefs-data` bucket, so restarting lakeFS with the same configuration brings the repository back. To delete everything, remove the bucket: + +```bash +rc rb rustfs/lakefs-data --force +``` + +## Troubleshooting + +### `failed to create repository: failed to access storage` with a DNS error + +lakeFS is using virtual-hosted addressing. Set `force_path_style: true` under `blockstore.s3` — in the 1.x configuration schema the older `path_style` key is rejected at startup. + +### `missing required keys: [auth.encrypt.secret_key]` + +lakeFS 1.x requires an encryption key for the local database. Add the `auth.encrypt.secret_key` block as shown in the configuration above. + +### `mkdir /lakefs: permission denied` + +The container runs as a non-root user and cannot create the database directory. Run the container with `-u 0` for local testing, or mount a writable directory at the configured path. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional lakeFS operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [lakeFS S3 blockstore documentation](https://docs.lakefs.io/howto/using-s3.html) to configure the S3 gateway for tools that speak the S3 protocol. diff --git a/content/ja/developer/integration/big-data/meta.json b/content/ja/developer/integration/big-data/meta.json index 114ef68b..6ebd60a9 100644 --- a/content/ja/developer/integration/big-data/meta.json +++ b/content/ja/developer/integration/big-data/meta.json @@ -2,16 +2,18 @@ "title": "データ分析", "pages": [ "clickhouse", + "duckdb", + "doris", + "flink", + "hudi", "iceberg", - "pyiceberg", + "influxdb", + "lakefs", "milvus", "mlflow", "opendal", - "duckdb", - "doris", - "influxdb", + "pyiceberg", "spark", - "flink", "trino", "zeppelin" ] diff --git a/content/ja/developer/integration/index.md b/content/ja/developer/integration/index.md index 5e47330c..23dd692e 100644 --- a/content/ja/developer/integration/index.md +++ b/content/ja/developer/integration/index.md @@ -8,8 +8,9 @@ description: "RustFS をリバースプロキシ、バックアップツール ## Integration categories - [Reverse Proxy](./reverse-proxy/index.md) は Nginx、Traefik、Caddy、HAProxy を扱います。 -- [Backup](./backup/index.md) は Restic と Longhorn を扱います。 -- [データ分析](./big-data/index.md) は ClickHouse、Doris、Iceberg、Milvus、OpenDAL、Zeppelin などの分析システムを扱います。 +- [Backup](./backup/index.md) は Kopia、Longhorn、Restic、Velero を扱います。 +- [AI](./ai/index.md) は Ray などの AI プラットフォームを扱います。 +- [データ分析](./big-data/index.md) は ClickHouse、Doris、Hudi、Iceberg、lakeFS、Milvus、OpenDAL、Zeppelin などの分析システムを扱います。 - [オブザーバビリティ](./observability/index.md) は Fluentd、OpenObserve、OpenTelemetry、Thanos、Tempo などのテレメトリシステムを扱います。 - [その他](./others/index.md) はコミュニティ主導の Python 用 capo SDK を扱います。 - [コンテナレジストリ](./registry/index.md) は Harbor を扱います。 diff --git a/content/ja/developer/integration/meta.json b/content/ja/developer/integration/meta.json index b4b29480..a3c989a5 100644 --- a/content/ja/developer/integration/meta.json +++ b/content/ja/developer/integration/meta.json @@ -4,6 +4,7 @@ "reverse-proxy", "backup", "big-data", + "ai", "observability", "others", "registry", diff --git a/content/zh/developer/integration/ai/images/rustfs-ray-data.png b/content/zh/developer/integration/ai/images/rustfs-ray-data.png new file mode 100644 index 00000000..aef8b561 Binary files /dev/null and b/content/zh/developer/integration/ai/images/rustfs-ray-data.png differ diff --git a/content/zh/developer/integration/ai/index.md b/content/zh/developer/integration/ai/index.md new file mode 100644 index 00000000..00a55f9d --- /dev/null +++ b/content/zh/developer/integration/ai/index.md @@ -0,0 +1,12 @@ +--- +title: "AI" +description: "通过 S3 兼容的对象存储接口,将 AI 平台连接到 RustFS。" +--- + +将 **RustFS** 用作支持 S3 兼容端点的 AI 与机器学习平台的对象存储层。 + +## 平台 + +- [Ray](./ray.md) + +请使用专用的存储桶保存训练数据集与检查点,并为凭证仅授予所需桶操作的权限。 diff --git a/content/zh/developer/integration/ai/meta.json b/content/zh/developer/integration/ai/meta.json new file mode 100644 index 00000000..573fc520 --- /dev/null +++ b/content/zh/developer/integration/ai/meta.json @@ -0,0 +1,6 @@ +{ + "title": "AI", + "pages": [ + "ray" + ] +} diff --git a/content/zh/developer/integration/ai/ray.md b/content/zh/developer/integration/ai/ray.md new file mode 100644 index 00000000..c08158db --- /dev/null +++ b/content/zh/developer/integration/ai/ray.md @@ -0,0 +1,103 @@ +--- +title: "Ray" +description: "使用 Ray Data 以 RustFS 作为 S3 兼容存储进行数据集读写。" +--- + +本指南通过 Ray Data 的 S3 文件系统支持,将分布式 AI 与 Python 计算框架 [Ray](https://github.com/ray-project/ray) 连接到 **RustFS**。你将在官方镜像内运行一个 Ray 作业,把数据集以 Parquet 写入 RustFS 存储桶,读回并验证对象。整个流程使用 `rayproject/ray:2.44.0-py311`(Ray 2.44,pyarrow 文件系统)和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Job["Ray job"] -->|"ray.data"| DS["Dataset"] + DS -->|"Parquet files"| RustFS["RustFS :9000"] +``` + +Ray Data 通过 pyarrow 的 `S3FileSystem` 读写数据集。传入为 RustFS 配置的 `S3FileSystem` 后,所有数据集操作——Parquet、CSV、JSON——都会指向该存储桶。 + +## 1. 创建作业文件 + +创建脚本,并替换全部连接占位符: + +```python title="ray_s3.py" +import ray +ray.init(ignore_reinit_error=True) + +import pandas as pd +from pyarrow.fs import S3FileSystem + +fs = S3FileSystem( + endpoint_override="http://:9000", + access_key="", + secret_key="", + region="us-east-1", +) + +df = pd.DataFrame({"id": range(5), "value": [x * 1.5 for x in range(5)]}) +ds = ray.data.from_pandas(df) +ds.write_parquet("my-bucket/ray-demo/events/", filesystem=fs) + +back = ray.data.read_parquet("my-bucket/ray-demo/events/", filesystem=fs).take_all() +print("rows:", len(back)) +print("sample:", back[0]) +ray.shutdown() +``` + +`endpoint_override` 接收带协议的完整端点 URL。pyarrow 的 `S3FileSystem` 对自定义端点自动使用 path-style 请求,无需额外参数。同一个文件系统对象也可用于 `write_csv`、`read_json` 等其他 Ray Data 方法。 + +## 2. 运行作业 + +在与 RustFS 相同的 Docker 网络中的 Ray 镜像内运行脚本: + +```bash +docker run --rm --network oo-rustfs_default \ + -v "$PWD/ray_s3.py":/tmp/ray_s3.py \ + rayproject/ray:2.44.0-py311 python /tmp/ray_s3.py +``` + +```text +rows: 5 +sample: {'id': 0, 'value': 0.0} +``` + +## 3. 在 RustFS 中验证对象 + +列出数据集前缀: + +```bash +rc ls rustfs/my-bucket/ray-demo/ -r +``` + +Ray Data 把数据集写成一个 Parquet 块: + +```text +ray-demo/events/0_000000_000000.parquet +``` + +![RustFS 控制台中存储的 Ray 数据集文件](./images/rustfs-ray-data.png) + +## 4. 停止或重置 + +Ray Data 自身不保存状态。删除演示数据集: + +```bash +rc rm rustfs/my-bucket/ray-demo/ --recursive --force +``` + +## 故障排查 + +### `Unable to connect to endpoint` 或超时 + +确认 `endpoint_override` 带协议且 Ray 容器可达。Compose 网络内主机名为 `rustfs`;宿主机上使用 `http://localhost:9000`。 + +### 写入时报 `Access Denied` + +确认访问密钥和秘密密钥直接传给了 `S3FileSystem`——这条代码路径不会读取容器的 AWS 环境变量。 + +## 后续步骤 + +- 在采用更多 Ray 操作之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Ray Data 文档](https://docs.ray.io/en/latest/data/data.html)在同一存储桶上串联转换、训练摄取与检查点流程。 diff --git a/content/zh/developer/integration/backup/images/rustfs-kopia-repo.png b/content/zh/developer/integration/backup/images/rustfs-kopia-repo.png new file mode 100644 index 00000000..a20e8f26 Binary files /dev/null and b/content/zh/developer/integration/backup/images/rustfs-kopia-repo.png differ diff --git a/content/zh/developer/integration/backup/images/rustfs-velero-backups.png b/content/zh/developer/integration/backup/images/rustfs-velero-backups.png new file mode 100644 index 00000000..569a2d6e Binary files /dev/null and b/content/zh/developer/integration/backup/images/rustfs-velero-backups.png differ diff --git a/content/zh/developer/integration/backup/index.md b/content/zh/developer/integration/backup/index.md index f56e073d..4b877961 100644 --- a/content/zh/developer/integration/backup/index.md +++ b/content/zh/developer/integration/backup/index.md @@ -8,6 +8,8 @@ description: "通过兼容 S3 的对象存储接口将备份工具连接到 Rust ## 系统 - [Restic](./restic.md) +- [Velero](./velero.md) +- [Kopia](./kopia.md) - [Longhorn](./longhorn.md) 请将备份作业放在专用的存储桶和前缀中,并使用仅限所需存储桶操作的凭据。 \ No newline at end of file diff --git a/content/zh/developer/integration/backup/kopia.md b/content/zh/developer/integration/backup/kopia.md new file mode 100644 index 00000000..18d03634 --- /dev/null +++ b/content/zh/developer/integration/backup/kopia.md @@ -0,0 +1,125 @@ +--- +title: "Kopia" +description: "使用 Kopia 的 S3 仓库后端把文件备份到 RustFS。" +--- + +本指南将开源备份恢复工具 [Kopia](https://github.com/kopia/kopia) 的 S3 仓库连接到 **RustFS**。你将在 RustFS 存储桶中创建仓库,对一个目录做快照,恢复到空目录,并比对校验和。整个流程使用 `kopia/kopia:0.18.1` 和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Source["Source files"] -->|snapshot| Kopia["Kopia"] + Kopia -->|"encrypted blocks"| RustFS["RustFS :9000"] + Kopia -->|restore| Restore["Restored files"] +``` + +Kopia 把仓库格式文件以及去重加密的内容块存进存储桶。恢复时读回内容块并重组原始文件,因此每个恢复文件的校验和必须与源文件一致。 + +## 1. 创建仓库 + +先创建存储桶,然后在其中初始化 Kopia 仓库。端点是纯 `host:port` 形式(不带协议);`--disable-tls` 让客户端使用纯 HTTP: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/kopia-backups + +docker run --rm --network oo-rustfs_default kopia/kopia:0.18.1 repository create s3 \ + --bucket kopia-backups \ + --access-key \ + --secret-access-key \ + --endpoint :9000 \ + --region us-east-1 \ + --disable-tls \ + --password \ + --override-username demo --override-hostname workstation +``` + +Kopia 在报告成功之前会通过 S3 API 读写来校验存储提供方。 + +## 2. 连接、快照、恢复 + +在保存 `repository.config` 的目录(由上一步生成)中执行以下命令。Kopia 从该文件读取连接设置,S3 参数只需提供一次: + +```bash +export KOPIA_PASSWORD= +export KOPIA_CONFIG_PATH=/config/repository.config + +alias kopia='docker run --rm --network oo-rustfs_default \ + -e KOPIA_PASSWORD -e KOPIA_CONFIG_PATH \ + -v "$PWD/config:/config" \ + -v "$PWD/source:/source:ro" \ + -v "$PWD/restore:/restore" kopia/kopia:0.18.1' + +kopia repository connect s3 \ + --bucket kopia-backups \ + --access-key \ + --secret-access-key \ + --endpoint :9000 \ + --region us-east-1 --disable-tls \ + --override-username demo --override-hostname workstation + +kopia snapshot create /source +kopia snapshot list +``` + +把快照恢复到空目录,并与源文件比对校验和。快照 ID 是 `snapshot list` 输出的 `ka...` 标识: + +```bash +kopia restore /restore + +sha256sum source/blob.bin restore/blob.bin +``` + +```text +bec4530e2798465b... source/blob.bin +bec4530e2798465b... restore/blob.bin +``` + +## 3. 在 RustFS 中验证对象 + +列出存储桶: + +```bash +rc ls rustfs/kopia-backups/ -r +``` + +输出包含仓库格式文件以及快照写入的打包内容块: + +```text +kopia.blobcfg +kopia.repository +p0000.../... +``` + +![RustFS 控制台中存储的 Kopia 仓库块](./images/rustfs-kopia-repo.png) + +## 4. 停止或重置 + +Kopia 是客户端工具,自身不保存运行状态。删除仓库和全部快照: + +```bash +rc rb rustfs/kopia-backups --force +``` + +## 故障排查 + +### `Endpoint url cannot have fully qualified paths` + +端点必须是纯 `host:port` 值,不带协议和路径——Kopia 自己构造对象 URL。 + +### `server gave HTTP response to HTTPS client` + +未加 `--disable-tls` 时 Kopia 使用 HTTPS。无 TLS 的 RustFS 需要在 `repository create s3` 和 `repository connect s3` 上都加 `--disable-tls`。 + +### `can't connect to storage` 并伴随 DNS 错误 + +S3 客户端正在使用 virtual-hosted 寻址。保持端点为纯 `host:port` 形式,Kopia 对这种端点使用 path-style 请求。 + +## 后续步骤 + +- 在采用更多 Kopia 操作之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Kopia 仓库文档](https://kopia.io/docs/repositories/)添加策略、保留与计划快照。 diff --git a/content/zh/developer/integration/backup/meta.json b/content/zh/developer/integration/backup/meta.json index 4511320b..0fb87d00 100644 --- a/content/zh/developer/integration/backup/meta.json +++ b/content/zh/developer/integration/backup/meta.json @@ -1,7 +1,9 @@ { "title": "备份", "pages": [ + "kopia", + "longhorn", "restic", - "longhorn" + "velero" ] } diff --git a/content/zh/developer/integration/backup/velero.md b/content/zh/developer/integration/backup/velero.md new file mode 100644 index 00000000..86fcb5b3 --- /dev/null +++ b/content/zh/developer/integration/backup/velero.md @@ -0,0 +1,153 @@ +--- +title: "Velero" +description: "使用 Velero 的 AWS 对象存储提供方把 Kubernetes 集群资源备份到 RustFS。" +--- + +本指南将 Kubernetes 备份恢复工具 [Velero](https://github.com/vmware-tanzu/velero) 连接到 **RustFS** 作为其对象存储后端。你将在 Kubernetes 集群中安装指向 RustFS 备份位置的 Velero,备份集群资源,删除后从存储桶恢复,并验证恢复结果。整个流程使用 Velero CLI 1.16.2、`velero-plugin-for-aws:v1.12.2`(k3s,Kubernetes 1.30)和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要一个可用 `kubectl` 访问的 Kubernetes 集群和 Velero CLI。本指南用于集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + K8s["Kubernetes cluster"] -->|"resources"| V["Velero"] + V -->|"backups + logs"| RustFS["RustFS :9000"] +``` + +Velero 把 Kubernetes 资源(以及可选的 Pod 卷数据)序列化为 gzip 归档,存到桶内 `backups/<名称>/` 前缀下。恢复时下载这些归档并在集群中重建资源。 + +## 1. 创建备份存储桶 + +先创建专用存储桶——Velero 不会创建桶: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/velero-backups +``` + +## 2. 安装 Velero + +把凭证写入文件并安装 Velero 服务端。`s3Url` 必须是集群 Pod 自己能访问的端点——请使用节点 IP 或内部地址,而不是工作站上的端口转发: + +```bash +cat > velero-creds < +aws_secret_access_key= +EOF + +kubectl create namespace velero +kubectl create secret generic cloud-credentials \ + --namespace velero --from-file=cloud=velero-creds + +velero install \ + --provider aws \ + --plugins velero/velero-plugin-for-aws:v1.12.2 \ + --bucket velero-backups \ + --backup-location-config region=us-east-1,s3ForcePathStyle="true",s3Url=http://:9000 \ + --secret-file velero-creds \ + --use-volume-snapshots=false +``` + +等待部署就绪、备份位置变为 `Available`: + +```bash +kubectl -n velero get pods +velero backup-location get +``` + +```text +NAME PROVIDER BUCKET/PREFIX PHASE LAST VALIDATED ACCESS MODE DEFAULT +default aws velero-backups Available 2026-09-22 ... ReadWrite true +``` + +## 3. 备份集群资源 + +创建两个演示资源并备份整个默认作用域: + +```bash +kubectl create configmap demo-cm --from-literal=key=rustfs-velero-demo +kubectl create deployment nginx --image=nginx:1.27 + +velero backup create demo-backup --wait +velero backup get +``` + +备份会把资源归档和日志上传到 RustFS。 + +## 4. 在 RustFS 中验证备份 + +列出存储桶: + +```bash +rc ls rustfs/velero-backups/ -r +``` + +备份归档集存放在 `backups/` 前缀下: + +```text +backups/demo-backup/demo-backup-resources.json.gz +backups/demo-backup/demo-backup-logs.gz +backups/demo-backup/demo-backup-itemoperations.json.gz +``` + +![RustFS 控制台中存储的 Velero 备份](./images/rustfs-velero-backups.png) + +## 5. 恢复并验证 + +删除演示资源后从备份恢复。恢复过程从 RustFS 读取归档并重建资源: + +```bash +kubectl delete configmap demo-cm +kubectl delete deployment nginx + +velero restore create --from-backup demo-backup --wait +``` + +确认资源已按原内容恢复: + +```bash +kubectl get configmap demo-cm -o jsonpath="{.data.key}" +kubectl get deployment nginx +``` + +```text +rustfs-velero-demo +NAME READY UP-TO-DATE AVAILABLE AGE +nginx 1/1 1 1 10s +``` + +## 6. 停止或重置 + +备份保留在 `velero-backups` 存储桶中,运行 Velero 并指向同一桶的任何集群都能恢复它们。从 RustFS 删除演示备份: + +```bash +velero backup delete demo-backup --confirm +rc rm rustfs/velero-backups/backups/ --recursive --force +``` + +## 故障排查 + +### 备份位置始终 `Unavailable` + +Velero Pod 无法访问 `s3Url`。Docker 自定义网桥网段对 Kubernetes Pod 不可路由——使用 Docker 宿主网桥 IP(单机集群为 `http://172.17.0.1:9000`)或节点地址。修补位置后 Velero 会重新校验: + +```bash +kubectl -n velero patch backupstoragelocation default --type merge -p \ + '{"spec":{"config":{"s3Url":"http://172.17.0.1:9000"}}}' +``` + +### 备份 `PartiallyFailed` 并报 daemonset 错误 + +`daemonset pod not found in running state` 错误表示 node-agent Pod(用于 Pod 卷备份)未运行。它不影响集群资源归档。等待 agent 就绪,或仅在 agent 运行时使用 `--default-volumes-to-fs-backup`。 + +### 安装后立即 `FailedValidation` + +第一次备份可能发生在位置完成校验之前。等 `velero backup-location get` 显示 `Available` 后再创建备份。 + +## 后续步骤 + +- 在采用更多 Velero 操作之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Velero 文档](https://velero.io/docs/main/)添加计划备份、卷快照与集群迁移流程。 diff --git a/content/zh/developer/integration/big-data/hudi.md b/content/zh/developer/integration/big-data/hudi.md new file mode 100644 index 00000000..1be2cea8 --- /dev/null +++ b/content/zh/developer/integration/big-data/hudi.md @@ -0,0 +1,109 @@ +--- +title: "Apache Hudi" +description: "通过 Spark 和 s3a 连接器把 Apache Hudi 表写入 RustFS。" +--- + +本指南将事务性数据湖平台 [Apache Hudi](https://github.com/apache/hudi) 连接到 **RustFS** 作为其写时复制表存储。你将运行带 Hudi bundle 的 Spark,向 RustFS 存储桶内的 `s3a://` 路径写入一张表,读回并验证表文件。整个流程使用 `apache/spark:3.5.6`、`hudi-spark3.5-bundle_2.12:0.15.0`、`hadoop-aws:3.3.4` 和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Job["Spark + Hudi"] -->|"commit + parquet"| RustFS["RustFS :9000"] +``` + +Hudi 把 `.hoodie/` 时间线、commit 文件和 Parquet 数据块存到桶内的表路径下。读取通过时间线解析表的最新快照,因此每次写入和查询都会经过 S3 API。 + +## 1. 运行 Spark 写入 + +Hudi 要求 Kryo 序列化器,s3a 凭证以 Hadoop 属性提供。下面的 `spark-shell` 会话在 `my-bucket` 中创建一张非分区表——请替换全部连接占位符: + +```scala +import org.apache.spark.sql.SaveMode + +val df = Seq((1, "a"), (2, "b"), (3, "c")).toDF("id", "name") +df.write.format("hudi") + .option("hoodie.table.name", "events") + .option("hoodie.datasource.write.recordkey.field", "id") + .option("hoodie.datasource.write.precombine.field", "name") + .option("hoodie.datasource.write.partitionpath.field", "") + .mode(SaveMode.Overwrite) + .save("s3a://my-bucket/hudi-demo/events") + +val back = spark.read.format("hudi").load("s3a://my-bucket/hudi-demo/events") +back.select("id", "name").show() +``` + +```bash +docker run --rm --network oo-rustfs_default \ + -v "$PWD/hudi_test.scala":/tmp/hudi_test.scala \ + apache/spark:3.5.6 /opt/spark/bin/spark-shell \ + --packages org.apache.hudi:hudi-spark3.5-bundle_2.12:0.15.0,org.apache.hadoop:hadoop-aws:3.3.4 \ + --conf spark.serializer=org.apache.spark.serializer.KryoSerializer \ + --conf spark.hadoop.fs.s3a.endpoint=http://rustfs:9000 \ + --conf spark.hadoop.fs.s3a.path.style.access=true \ + --conf spark.hadoop.fs.s3a.access.key= \ + --conf spark.hadoop.fs.s3a.secret.key= \ + --conf spark.hadoop.fs.s3a.region=us-east-1 \ + --conf spark.jars.ivy=/tmp/.ivy2 \ + --conf spark.sql.shuffle.partitions=2 \ + -i /tmp/hudi_test.scala +``` + +`--packages` 参数会在首次运行时下载 Hudi bundle 和 s3a 连接器。`spark.jars.ivy` 设置用于避免容器内的 Ivy 缓存权限错误。读回的 `show()` 输出三行数据: + +```text ++---+----+ +| 2| b| +| 3| c| +| 1| a| ++---+----+ +``` + +## 2. 在 RustFS 中验证对象 + +列出表前缀: + +```bash +rc ls rustfs/my-bucket/hudi-demo/ -r +``` + +表路径下出现 Hudi 表布局——`.hoodie/` 时间线及 commit 文件,加上 Parquet 数据块: + +```text +hudi-demo/events/.hoodie/20260922015803291.commit +hudi-demo/events/.hoodie/20260922015803291.commit.requested +hudi-demo/events/// +``` + +![RustFS 控制台中存储的 Hudi 表文件](./images/rustfs-hudi-table.png) + +## 3. 停止或重置 + +Spark 是一次性客户端,不保存状态。删除演示表: + +```bash +rc rm rustfs/my-bucket/hudi-demo/ --recursive --force +``` + +## 故障排查 + +### `hoodie only support org.apache.spark.serializer.KryoSerializer as spark.serializer` + +Hudi 拒绝 Spark 默认序列化器。按上文传入 `--conf spark.serializer=org.apache.spark.serializer.KryoSerializer`。 + +### `Partition-path field has to be non-empty` 或 keygenerator 类错误 + +非分区表保持默认 key generator(不要设置 `hoodie.datasource.write.keygenerator.class`),并把 `hoodie.datasource.write.partitionpath.field` 设为空字符串。 + +### Hudi 写入时报 `NoSuchMethodError` 或 `ClassNotFoundException` + +Spark 与 Hudi bundle 版本必须匹配:`hudi-spark3.5-bundle_2.12` 搭配 Spark 3.5.x,内置 Hadoop 客户端需兼容 `hadoop-aws:3.3.4`。 + +## 后续步骤 + +- 在采用更多 Hudi 操作之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Hudi Spark 指南](https://hudi.apache.org/docs/quick-start-guide/)在同一张表上添加 upsert、compaction 与查询集成。 diff --git a/content/zh/developer/integration/big-data/images/rustfs-hudi-table.png b/content/zh/developer/integration/big-data/images/rustfs-hudi-table.png new file mode 100644 index 00000000..0db15da9 Binary files /dev/null and b/content/zh/developer/integration/big-data/images/rustfs-hudi-table.png differ diff --git a/content/zh/developer/integration/big-data/images/rustfs-lakefs-repo.png b/content/zh/developer/integration/big-data/images/rustfs-lakefs-repo.png new file mode 100644 index 00000000..44e4e567 Binary files /dev/null and b/content/zh/developer/integration/big-data/images/rustfs-lakefs-repo.png differ diff --git a/content/zh/developer/integration/big-data/index.md b/content/zh/developer/integration/big-data/index.md index ecce2ee7..fb0d70cf 100644 --- a/content/zh/developer/integration/big-data/index.md +++ b/content/zh/developer/integration/big-data/index.md @@ -8,6 +8,7 @@ description: "通过 S3 兼容的对象存储接口将数据分析系统连接 ## 系统 - [ClickHouse](./clickhouse.md) +- [Hudi](./hudi.md) - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) - [Milvus](./milvus.md) @@ -15,6 +16,7 @@ description: "通过 S3 兼容的对象存储接口将数据分析系统连接 - [OpenDAL](./opendal.md) - [DuckDB](./duckdb.md) - [Doris](./doris.md) +- [lakeFS](./lakefs.md) - [InfluxDB](./influxdb.md) - [Spark](./spark.md) - [Flink](./flink.md) diff --git a/content/zh/developer/integration/big-data/lakefs.md b/content/zh/developer/integration/big-data/lakefs.md new file mode 100644 index 00000000..ac4ea634 --- /dev/null +++ b/content/zh/developer/integration/big-data/lakefs.md @@ -0,0 +1,160 @@ +--- +title: "lakeFS" +description: "以 RustFS 作为 lakeFS 的 S3 blockstore,构建带版本管理的数据湖。" +--- + +本指南将 Git 式数据湖版本管理层 [lakeFS](https://github.com/treeverse/lakeFS) 连接到 **RustFS** 作为其 S3 blockstore。你将启动 lakeFS,创建存储命名空间指向 RustFS 存储桶的仓库,提交一个对象,并验证 lakeFS 的元数据和数据文件存放在 RustFS 中。整个流程使用 `treeverse/lakefs:1.58.0` 和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Client["lakectl / API"] -->|HTTP| LakeFS["lakeFS :8000"] + LakeFS -->|"metadata + data"| RustFS["RustFS :9000"] +``` + +lakeFS 把仓库元数据和已提交的数据文件按存储命名空间存入 blockstore。桶内是一个 `repo/` 前缀,包含 `_lakefs/` 元数据和 `data/` 对象,全部通过 S3 API 读写。 + +## 1. 创建项目文件 + +先创建存储桶,再创建 lakeFS 配置: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/lakefs-data +``` + +```yaml title="config.yaml" +database: + type: local + local: + path: /lakefs/data + +blockstore: + type: s3 + s3: + endpoint: http://rustfs:9000 + region: us-east-1 + force_path_style: true + +auth: + encrypt: + secret_key: change-me-to-a-random-string + +gateways: + s3: + domain_name: rustfs-gateway.local + +logging: + format: text + level: info +``` + +`force_path_style: true` 是必需的——缺少它 lakeFS 会把 `lakefs-data.<主机名>` 当作主机名,所有请求都会因 DNS 错误失败。凭证通过标准的 `AWS_ACCESS_KEY_ID` 和 `AWS_SECRET_ACCESS_KEY` 环境变量提供。 + +创建服务凭证与初始管理员用户的环境文件,并替换占位符: + +```ini title=".env" +AWS_ACCESS_KEY_ID= +AWS_SECRET_ACCESS_KEY= +LAKEFS_INSTALLATION_USER_NAME=admin +LAKEFS_INSTALLATION_ACCESS_KEY_ID= +LAKEFS_INSTALLATION_SECRET_ACCESS_KEY= +LAKEFS_STATS_ENABLED=false +``` + +## 2. 启动 lakeFS + +```bash +docker run -d --name lakefs --network oo-rustfs_default \ + -p 8000:8000 \ + -v "$PWD/config.yaml":/etc/lakefs/config.yaml:ro \ + --env-file .env \ + treeverse/lakefs:1.58.0 run +``` + +等待 `http://localhost:8000/api/healthcheck` 返回 `200`,然后创建存储命名空间位于桶内的仓库: + +```bash +curl -s -u : \ + -X POST http://localhost:8000/api/v1/repositories \ + -H "Content-Type: application/json" \ + -d '{"name": "rustfs-demo", "storage_namespace": "s3://lakefs-data/repo"}' +``` + +## 3. 提交一个对象 + +向 `main` 分支上传文件并提交: + +```bash +echo "hello from lakefs on rustfs" > hello.txt + +curl -s -u : \ + -X POST "http://localhost:8000/api/v1/repositories/rustfs-demo/branches/main/objects?path=hello.txt" \ + --data-binary @hello.txt + +curl -s -u : \ + -X POST "http://localhost:8000/api/v1/repositories/rustfs-demo/branches/main/commits" \ + -H "Content-Type: application/json" \ + -d '{"message": "add hello"}' +``` + +通过分支 ref 读回对象——响应体就是已提交的内容: + +```bash +curl -s -u : \ + "http://localhost:8000/api/v1/repositories/rustfs-demo/refs/main/objects?path=hello.txt" +``` + +## 4. 在 RustFS 中验证对象 + +列出存储桶: + +```bash +rc ls rustfs/lakefs-data/ -r +``` + +输出显示仓库前缀下的 lakeFS 元数据对象和已提交数据文件: + +```text +repo/_lakefs/19b2b26e37cb20fc6763c527f88eb5151891b04a2c8c9ddd32870c5c3f353281 +repo/data/fueia10jdra000e1c480/daots7ojdra000e1c490 +``` + +![RustFS 控制台中存储的 lakeFS 对象](./images/rustfs-lakefs-repo.png) + +## 5. 停止或重置部署 + +停止 lakeFS 并保留数据: + +```bash +docker rm -f lakefs +``` + +仓库元数据和数据保留在 `lakefs-data` 存储桶中,以相同配置重启 lakeFS 即可恢复仓库。若要删除全部内容,请移除存储桶: + +```bash +rc rb rustfs/lakefs-data --force +``` + +## 故障排查 + +### `failed to create repository: failed to access storage` 并伴随 DNS 错误 + +lakeFS 正在使用 virtual-hosted 寻址。在 `blockstore.s3` 下设置 `force_path_style: true`——1.x 配置 schema 会拒绝旧的 `path_style` 键。 + +### `missing required keys: [auth.encrypt.secret_key]` + +lakeFS 1.x 要求本地数据库的加密密钥。按上文配置添加 `auth.encrypt.secret_key` 段。 + +### `mkdir /lakefs: permission denied` + +容器以非 root 用户运行,无法创建数据库目录。本地测试可加 `-u 0` 运行,或在配置路径上挂载可写目录。 + +## 后续步骤 + +- 在采用更多 lakeFS 操作之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [lakeFS S3 blockstore 文档](https://docs.lakefs.io/howto/using-s3.html)配置 S3 网关,让讲 S3 协议的工具直接接入。 diff --git a/content/zh/developer/integration/big-data/meta.json b/content/zh/developer/integration/big-data/meta.json index ecf05567..f56c8131 100644 --- a/content/zh/developer/integration/big-data/meta.json +++ b/content/zh/developer/integration/big-data/meta.json @@ -2,16 +2,18 @@ "title": "数据分析", "pages": [ "clickhouse", + "duckdb", + "doris", + "flink", + "hudi", "iceberg", - "pyiceberg", + "influxdb", + "lakefs", "milvus", "mlflow", "opendal", - "duckdb", - "doris", - "influxdb", + "pyiceberg", "spark", - "flink", "trino", "zeppelin" ] diff --git a/content/zh/developer/integration/index.md b/content/zh/developer/integration/index.md index f455c86d..24b32f8d 100644 --- a/content/zh/developer/integration/index.md +++ b/content/zh/developer/integration/index.md @@ -8,8 +8,9 @@ description: "将 RustFS 与反向代理、备份工具、数据分析系统、 ## 集成类别 - [反向代理](./reverse-proxy/index.md)涵盖 Nginx、Traefik、Caddy 和 HAProxy。 -- [备份](./backup/index.md)涵盖 Restic 和 Longhorn。 -- [数据分析](./big-data/index.md)涵盖 ClickHouse、Doris、Iceberg、Milvus、OpenDAL 和 Zeppelin 等数据分析系统。 +- [备份](./backup/index.md)涵盖 Kopia、Longhorn、Restic 和 Velero。 +- [AI](./ai/index.md)涵盖 Ray 等 AI 平台。 +- [数据分析](./big-data/index.md)涵盖 ClickHouse、Doris、Hudi、Iceberg、lakeFS、Milvus、OpenDAL 和 Zeppelin 等数据分析系统。 - [可观测性](./observability/index.md)涵盖 Fluentd、OpenObserve、OpenTelemetry、Thanos 和 Tempo 等遥测系统。 - [其他](./others/index.md)涵盖社区驱动的 Python capo SDK。 - [镜像仓库](./registry/index.md)涵盖 Harbor。 diff --git a/content/zh/developer/integration/meta.json b/content/zh/developer/integration/meta.json index b4b29480..a3c989a5 100644 --- a/content/zh/developer/integration/meta.json +++ b/content/zh/developer/integration/meta.json @@ -4,6 +4,7 @@ "reverse-proxy", "backup", "big-data", + "ai", "observability", "others", "registry",