diff --git a/content/de/developer/integration/backup/images/rustfs-longhorn-backups.png b/content/de/developer/integration/backup/images/rustfs-longhorn-backups.png new file mode 100644 index 00000000..899b1c8c Binary files /dev/null and b/content/de/developer/integration/backup/images/rustfs-longhorn-backups.png differ diff --git a/content/de/developer/integration/backup/index.md b/content/de/developer/integration/backup/index.md index eabb9aad..de236c2a 100644 --- a/content/de/developer/integration/backup/index.md +++ b/content/de/developer/integration/backup/index.md @@ -8,5 +8,6 @@ Verwenden Sie **RustFS** als Objektspeicher-Backend für Backup-Tools, die Repos ## Systeme - [Restic](./restic.md) +- [Longhorn](./longhorn.md) Halten Sie Backup-Jobs in einem eigenen Bucket und Präfix und verwenden Sie Anmeldedaten, die auf die erforderlichen Bucket-Operationen beschränkt sind. \ No newline at end of file diff --git a/content/de/developer/integration/backup/longhorn.md b/content/de/developer/integration/backup/longhorn.md new file mode 100644 index 00000000..497f43cd --- /dev/null +++ b/content/de/developer/integration/backup/longhorn.md @@ -0,0 +1,293 @@ +--- +title: "Longhorn" +description: "Configure Longhorn to store Kubernetes volume backups in RustFS through its S3 backup target." +--- + +This guide connects [Longhorn](https://github.com/longhorn/longhorn) — the distributed block storage system for Kubernetes — to **RustFS** as its S3 backup target. You will configure the backup target, back up a volume that contains data, delete the volume, restore it from RustFS, and verify the data. The workflow was verified with Longhorn 1.9.0 on k3s (Kubernetes 1.30) and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need a Kubernetes cluster with Longhorn installed and `kubectl` access to it. This guide is intended for integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + App["Workload pod"] -->|"writes"| Vol["Longhorn volume"] + Vol -->|"snapshot"| Backup["Backup engine"] + Backup -->|"blocks + config"| RustFS["RustFS :9000"] +``` + +Longhorn stores backups in the `backupstore/volumes/` prefix of the target bucket as content-addressed blocks plus a small volume configuration file. Restores read those objects back into a new volume on any node that can reach RustFS. + +## 1. Create the backup bucket + +Create a dedicated bucket with the [`rc` client](https://github.com/rustfs/cli), using your own endpoint and credentials: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/longhorn-backups +``` + +Longhorn does not create the bucket, so this step must complete before the first backup. + +## 2. Configure the backup target + +Store the RustFS credentials in a secret in the `longhorn-system` namespace. `AWS_ENDPOINTS` must be an endpoint every node can reach — use the node IP or an internal load balancer address, not a port forward from your workstation: + +```bash +kubectl -n longhorn-system create secret generic rustfs-s3-secret \ + --from-literal=AWS_ACCESS_KEY_ID= \ + --from-literal=AWS_SECRET_ACCESS_KEY= \ + --from-literal=AWS_ENDPOINTS=http://:9000 +``` + +Longhorn 1.9 manages backup targets through the `BackupTarget` resource. Patch the `default` target with the RustFS bucket: + +```bash +kubectl -n longhorn-system patch backupTarget default --type merge -p ' +spec: + backupTargetURL: s3://longhorn-backups@us-east-1/ + credentialSecret: rustfs-s3-secret + pollInterval: 5m' +``` + +The `@us-east-1` segment is the region annotation in the S3 URL format; it does not need to match a real deployment region. + +Wait for the target to become available — it confirms that Longhorn reached the bucket through the secret: + +```bash +kubectl -n longhorn-system get backupTarget default +``` + +```text +NAME URL CREDENTIAL AVAILABLE LASTSYNCEDAT +default s3://longhorn-backups@us-east-1/ rustfs-s3-secret true 2026-09-21T08:36:55Z +``` + +## 3. Write data and create a backup + +Create a test volume with data in it. The following manifest creates a 1 GiB PVC and a pod that writes a marker file: + +```yaml title="demo.yaml" +apiVersion: v1 +kind: PersistentVolumeClaim +metadata: + name: demo-vol +spec: + accessModes: [ReadWriteOnce] + storageClassName: longhorn + resources: + requests: + storage: 1Gi +--- +apiVersion: v1 +kind: Pod +metadata: + name: demo-app +spec: + volumes: + - name: data + persistentVolumeClaim: + claimName: demo-vol + containers: + - name: app + image: busybox:1.36 + command: ["sh", "-c", "echo 'longhorn rustfs demo' > /data/hello.txt && sleep 3600"] + volumeMounts: + - name: data + mountPath: /data +``` + +Apply it and wait for the pod to run: + +```bash +kubectl apply -f demo.yaml +kubectl get pod demo-app +``` + +Create a snapshot and back it up. You can do this in the Longhorn UI (Volume → Snapshot → Backup) or declaratively: + +```bash +VOLUME=$(kubectl get pvc demo-vol -o jsonpath='{.spec.volumeName}') +kubectl -n longhorn-system apply -f - <" +EOF +``` + +The restore runs when the volume is first attached. Create a PVC bound to the restored volume through a static PV, then mount it: + +```bash +kubectl apply -f - < --type merge -p ' +spec: + disks: + default: + path: /var/lib/longhorn + allowScheduling: true' +``` + +### `failed to create backup ... missing input parameter` + +The backup ran before the default backup target was configured. Complete step 2, confirm `AVAILABLE` is `true`, and create the backup again. + +### The backup target never becomes available + +The secret must exist before the target syncs, and `AWS_ENDPOINTS` must be reachable from the nodes themselves. Check the `longhorn-manager` logs for S3 errors: + +```bash +kubectl -n longhorn-system logs -l app=longhorn-manager | grep -i s3 | tail +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional backup targets. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Longhorn backup documentation](https://longhorn.io/docs/1.9.0/backups-and-restore/) to configure recurring backup jobs and scheduled snapshots. diff --git a/content/de/developer/integration/backup/meta.json b/content/de/developer/integration/backup/meta.json index 9de69ffa..3ea20cbe 100644 --- a/content/de/developer/integration/backup/meta.json +++ b/content/de/developer/integration/backup/meta.json @@ -1,6 +1,7 @@ { "title": "Backup", "pages": [ - "restic" + "restic", + "longhorn" ] -} \ No newline at end of file +} diff --git a/content/de/developer/integration/big-data/images/rustfs-mlflow-artifacts.png b/content/de/developer/integration/big-data/images/rustfs-mlflow-artifacts.png new file mode 100644 index 00000000..545f5bf2 Binary files /dev/null and b/content/de/developer/integration/big-data/images/rustfs-mlflow-artifacts.png differ diff --git a/content/de/developer/integration/big-data/index.md b/content/de/developer/integration/big-data/index.md index 2ae0de3d..ef0b7cba 100644 --- a/content/de/developer/integration/big-data/index.md +++ b/content/de/developer/integration/big-data/index.md @@ -10,6 +10,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) - [Milvus](./milvus.md) +- [MLflow](./mlflow.md) - [DuckDB](./duckdb.md) - [InfluxDB](./influxdb.md) - [Spark](./spark.md) diff --git a/content/de/developer/integration/big-data/meta.json b/content/de/developer/integration/big-data/meta.json index 1d6f178f..87e29f91 100644 --- a/content/de/developer/integration/big-data/meta.json +++ b/content/de/developer/integration/big-data/meta.json @@ -4,6 +4,7 @@ "iceberg", "pyiceberg", "milvus", + "mlflow", "duckdb", "influxdb", "spark", diff --git a/content/de/developer/integration/big-data/mlflow.md b/content/de/developer/integration/big-data/mlflow.md new file mode 100644 index 00000000..46d3205e --- /dev/null +++ b/content/de/developer/integration/big-data/mlflow.md @@ -0,0 +1,244 @@ +--- +title: "MLflow" +description: "Run MLflow with RustFS as the S3 artifact store for experiment tracking, deployed with Docker Compose." +--- + +This guide connects [MLflow](https://github.com/mlflow/mlflow) — the experiment tracking and model registry platform — to **RustFS** as its S3 artifact store. You will start the MLflow tracking server with Docker Compose, log parameters, metrics, and artifacts from a training run, and verify that the artifacts are stored in RustFS. The workflow was verified with `ghcr.io/mlflow/mlflow:v2.22.1` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with the Compose plugin. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["Training script"] -->|"runs + metrics"| Server["MLflow server :5000"] + Client -->|"artifacts"| RustFS["RustFS :9000"] + Server -->|"metadata"| DB["SQLite"] +``` + +The tracking server keeps experiment and run metadata in SQLite and stores artifacts — model files, plots, reports — directly in RustFS through the `s3://` artifact root. The client needs the same RustFS credentials because it uploads artifacts itself, using boto3 and the `MLFLOW_S3_ENDPOINT_URL` setting. + +## 1. Create the project files + +Create a working directory: + +```bash +mkdir rustfs-mlflow +cd rustfs-mlflow +``` + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +RUSTFS_ACCESS_KEY= +RUSTFS_SECRET_KEY= +MLFLOW_BUCKET=my-bucket +``` + +Use dedicated credentials for the artifact bucket. Do not commit `.env` to source control. + +Create the Compose file: + +```yaml title="compose.yaml" +services: + rustfs: + image: rustfs/rustfs-x86-musl:v2.3.1 + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + RUSTFS_VOLUMES: /data + RUSTFS_ADDRESS: ":9000" + RUSTFS_CONSOLE_ADDRESS: ":9001" + RUSTFS_CONSOLE_ENABLE: "true" + volumes: + - rustfs-data:/data + ports: + - "9000:9000" + - "9001:9001" + healthcheck: + test: ["CMD", "curl", "-sf", "http://127.0.0.1:9000/health"] + interval: 10s + timeout: 5s + retries: 6 + start_period: 10s + networks: + - mlflow + + create-bucket: + image: rustfs/rc:latest + depends_on: + rustfs: + condition: service_healthy + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + MLFLOW_BUCKET: ${MLFLOW_BUCKET} + entrypoint: + - /bin/sh + - -c + - | + /usr/bin/rc alias set rustfs http://rustfs:9000 "$${RUSTFS_ACCESS_KEY}" "$${RUSTFS_SECRET_KEY}" + /usr/bin/rc mb --ignore-existing rustfs/$${MLFLOW_BUCKET} + networks: + - mlflow + + mlflow: + image: ghcr.io/mlflow/mlflow:v2.22.1 + command: + - server + - --backend-store-uri + - sqlite:////mlflow/mlflow.db + - --default-artifact-root + - s3://${MLFLOW_BUCKET}/mlflow-artifacts + - --host + - 0.0.0.0 + - --port + - "5000" + environment: + AWS_ACCESS_KEY_ID: ${RUSTFS_ACCESS_KEY} + AWS_SECRET_ACCESS_KEY: ${RUSTFS_SECRET_KEY} + MLFLOW_S3_ENDPOINT_URL: http://rustfs:9000 + AWS_DEFAULT_REGION: us-east-1 + depends_on: + create-bucket: + condition: service_completed_successfully + ports: + - "5000:5000" + volumes: + - mlflow-db:/mlflow + networks: + - mlflow + +networks: + mlflow: + +volumes: + rustfs-data: + mlflow-db: +``` + +The `create-bucket` service must run before the server starts because MLflow does not create the bucket. `MLFLOW_S3_ENDPOINT_URL` points the server's boto3 client at RustFS with path-style addressing. The SQLite database lives on a volume so experiment metadata survives restarts; production deployments should use a managed database backend instead. + +## 2. Validate and start the deployment + +Resolve the Compose file before starting containers: + +```bash +docker compose config +``` + +Start the services and wait for the tracking server: + +```bash +docker compose up -d +docker compose ps +``` + +The `create-bucket` service should exit with code `0`, and the MLflow UI should answer on `http://localhost:5000`: + +```bash +curl -sf http://localhost:5000/ >/dev/null && echo ready +``` + +Open the RustFS Console at `http://localhost:9001` to watch artifacts land in `my-bucket` during the next step. + +## 3. Log a training run + +Create the client script — it runs inside the MLflow image, which ships every dependency: + +```bash title="train_demo.py" {12} +import mlflow + +mlflow.set_tracking_uri("http://localhost:5000") +mlflow.set_experiment("rustfs-demo") + +with mlflow.start_run(run_name="rustfs-verify") as run: + mlflow.log_params({"model": "demo-regressor", "alpha": 0.5}) + for step in range(3): + mlflow.log_metric("rmse", 0.9 - step * 0.2, step=step) + with open("model-summary.txt", "w") as f: + f.write("demo model trained against RustFS artifact store\n") + mlflow.log_artifact("model-summary.txt", artifact_path="reports") + print("run_id:", run.info.run_id) + print("artifact_uri:", run.info.artifact_uri) +``` + +Copy the script into the running container and execute it with the service environment: + +```bash +docker compose cp train_demo.py mlflow:/tmp/train_demo.py +docker compose exec -w /tmp mlflow python train_demo.py +``` + +The `artifact_uri` prints as `s3://my-bucket/mlflow-artifacts///artifacts` — the artifact is uploaded straight to RustFS by the client. + +## 4. Verify artifacts in RustFS + +Read the artifact back through the tracking server, then list the same object in RustFS. Save the following as `verify.py`, replace `` with the identifier printed in step 3, and run it the same way: + +```bash title="verify.py" {5} +import mlflow + +mlflow.set_tracking_uri("http://localhost:5000") +client = mlflow.MlflowClient() +print([a.path for a in client.list_artifacts("", "reports")]) +path = client.download_artifacts("", "reports/model-summary.txt") +print(open(path).read()) +``` + +The download reads the object from RustFS through the tracking server. Then confirm the objects in the bucket: + +```bash +docker compose exec rustfs /usr/bin/rc ls local/my-bucket/mlflow-artifacts/ -r +``` + +The output should include the artifact object: + +```text +mlflow-artifacts/1/a1aece9243504f2680a52fba0c32765f/artifacts/reports/model-summary.txt +``` + +![MLflow artifacts stored in the RustFS Console](./images/rustfs-mlflow-artifacts.png) + +Runs, parameters, and metrics survive a server restart because they are stored in SQLite, while the artifacts stay in RustFS: + +```bash +docker compose restart mlflow +docker compose exec -w /tmp mlflow python verify.py +``` + +The `list_artifacts` call succeeds again against the restarted server, reading from the same objects in RustFS. + +## 5. Stop or reset the deployment + +Stop the containers while keeping all data: + +```bash +docker compose down +``` + +The RustFS volume keeps the artifact objects and the MLflow volume keeps the metadata database. To delete everything, including the artifacts in RustFS, add `--volumes`. + +## Troubleshooting + +### `ModuleNotFoundError: No module named 'boto3'` when logging artifacts + +The client performing `log_artifact` needs boto3 because it uploads directly to S3. Install it in the environment that runs the training script, or run the script inside the MLflow image as shown in step 3. + +### `AccessDenied` or connection errors during artifact upload + +Confirm that `MLFLOW_S3_ENDPOINT_URL` is set in the client environment — without it, boto3 sends requests to real AWS S3. The endpoint must be reachable from the machine running the training script; use `http://localhost:9000` outside the Compose network and `http://rustfs:9000` inside it. + +### The server fails to start with a bucket error + +The artifact bucket must exist before the server starts. Check the `create-bucket` service logs: + +```bash +docker compose logs create-bucket +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional MLflow operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [MLflow documentation](https://mlflow.org/docs/latest/) to add a model registry or move the metadata store to a managed database. diff --git a/content/de/developer/integration/devops/gitea.md b/content/de/developer/integration/devops/gitea.md new file mode 100644 index 00000000..239b6ad2 --- /dev/null +++ b/content/de/developer/integration/devops/gitea.md @@ -0,0 +1,236 @@ +--- +title: "Gitea" +description: "Run Gitea with RustFS as the S3 storage backend for LFS objects and attachments, deployed with Docker Compose." +--- + +This guide connects [Gitea](https://github.com/go-gitea/gitea) — the self-hosted Git service — to **RustFS** through Gitea's `minio` storage type. You will start Gitea with Docker Compose, create a repository, push a Git LFS object, attach a file to an issue, and verify that both the LFS object and the attachment are stored in RustFS. The workflow was verified with `gitea/gitea:1.24.4` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with the Compose plugin and the `git` and `git-lfs` clients on your workstation. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Dev["Git + LFS client"] -->|"git push / git-lfs"| Gitea["Gitea :3000"] + Gitea -->|"LFS + attachments"| RustFS["RustFS :9000"] +``` + +Gitea stores the Git repository itself on its local disk, while the `minio` storage type routes large files — LFS objects, issue attachments, avatars, repository archives, packages, and Actions artifacts — to the `gitea-data` bucket in RustFS. Gitea creates the bucket on startup if it does not exist. + +## 1. Create the project files + +Create a working directory: + +```bash +mkdir rustfs-gitea +cd rustfs-gitea +``` + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +RUSTFS_ACCESS_KEY= +RUSTFS_SECRET_KEY= +``` + +Use dedicated credentials for the `gitea-data` bucket. Do not commit `.env` to source control. + +Create the Gitea configuration and replace both credential placeholders with the same values you set in `.env` — the global `minio` storage type applies to LFS, attachments, avatars, repository archives, packages, and Actions artifacts: + +```ini title="app.ini" +APP_NAME = RustFS Gitea +RUN_MODE = prod +WORK_PATH = /data/gitea + +[server] +DOMAIN = localhost +ROOT_URL = http://localhost:3000/ +HTTP_PORT = 3000 +LFS_START_SERVER = true + +[database] +DB_TYPE = sqlite3 +PATH = /data/gitea/gitea.db + +[storage] +STORAGE_TYPE = minio +MINIO_ENDPOINT = rustfs:9000 +MINIO_ACCESS_KEY_ID = +MINIO_SECRET_ACCESS_KEY = +MINIO_BUCKET = gitea-data +MINIO_LOCATION = us-east-1 +MINIO_USE_SSL = false + +[log] +MODE = console +LEVEL = info + +[security] +INSTALL_LOCK = true +SECRET_KEY = change-me-to-a-random-string +``` + +`LFS_START_SERVER` enables the Git LFS HTTP API. `MINIO_ENDPOINT` uses the Compose-internal hostname `rustfs`; the endpoint is reached with path-style requests by default. + +Create the Compose file: + +```yaml title="compose.yaml" +services: + rustfs: + image: rustfs/rustfs-x86-musl:v2.3.1 + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + RUSTFS_VOLUMES: /data + RUSTFS_ADDRESS: ":9000" + RUSTFS_CONSOLE_ADDRESS: ":9001" + RUSTFS_CONSOLE_ENABLE: "true" + volumes: + - rustfs-data:/data + ports: + - "9000:9000" + - "9001:9001" + healthcheck: + test: ["CMD", "curl", "-sf", "http://127.0.0.1:9000/health"] + interval: 10s + timeout: 5s + retries: 6 + start_period: 10s + networks: + - gitea + + gitea: + image: gitea/gitea:1.24.4 + depends_on: + rustfs: + condition: service_healthy + environment: + USER_UID: "1000" + USER_GID: "1000" + volumes: + - gitea-data:/data + - ./app.ini:/data/gitea/conf/app.ini:ro + ports: + - "3000:3000" + networks: + - gitea + +networks: + gitea: + +volumes: + rustfs-data: + gitea-data: +``` + +The volume-backed `/data` directory keeps the SQLite database and the Git repositories across container restarts, while LFS objects and attachments live in RustFS. + +## 2. Validate and start the deployment + +Resolve the Compose file before starting containers: + +```bash +docker compose config +``` + +Start the services: + +```bash +docker compose up -d +docker compose ps +``` + +Watch the Gitea log until every storage backend reports the Minio type: + +```bash +docker compose logs gitea | grep "Initialising" +``` + +The output should list `Attachment`, `Avatar`, `LFS`, and the remaining storage sections, each followed by a `Creating Minio storage at rustfs:9000:gitea-data` line. + +Open `http://localhost:3000` and create the administrator account, then open the RustFS Console at `http://localhost:9001` — the `gitea-data` bucket appears after the first storage operation. + +## 3. Push a Git LFS object + +Create a repository named `rustfs-demo` in the Gitea web UI, then push an LFS-tracked file from your workstation: + +```bash +mkdir lfs-demo && cd lfs-demo +git init +git config user.email you@example.com +git config user.name you +git lfs install +git lfs track "*.bin" +git add .gitattributes +dd if=/dev/urandom of=dataset.bin bs=1M count=8 +git add dataset.bin +git commit -m "add LFS dataset" +git remote add origin http://localhost:3000//rustfs-demo.git +git push origin main +``` + +`git push` uploads the LFS object through the Gitea LFS API, which writes it to RustFS. Clone the repository into a second directory and run `git lfs pull` — the downloaded `dataset.bin` must be byte-identical to the original: + +```bash +sha256sum dataset.bin +cd ../lfs-demo-clone && git lfs pull && sha256sum dataset.bin +``` + +Both checksums match because both clients read the object from RustFS. + +## 4. Attach a file to an issue + +Open the `rustfs-demo` repository, create an issue, and attach a small text file through the issue form. Gitea stores the upload as `attachments//` in the `gitea-data` bucket and serves downloads through `/attachments/`. + +## 5. Verify objects in RustFS + +List the bucket with the [`rc` client](https://github.com/rustfs/cli): + +```bash +docker compose exec rustfs /usr/bin/rc ls local/gitea-data/ -r +``` + +The output should include the LFS object under `lfs/` and the attachment under `attachments/`: + +```text +attachments/9/2/92fdd48d-531c-4cba-8b3f-4e2004a10fc7 +lfs/37/76/6ddfc07e803de58a69328db9a58a07cf7080ddde55c155a7531bc650a000 +``` + +The LFS object key is the SHA-256 content hash used by the Git LFS protocol. + +![Gitea LFS and attachment objects in the RustFS Console](./images/rustfs-gitea-objects.png) + +## 6. Stop or reset the deployment + +Stop the containers while keeping all data: + +```bash +docker compose down +``` + +The RustFS volume keeps the `gitea-data` bucket, so LFS objects and attachments survive a restart. To delete everything, including the objects in RustFS, add `--volumes`. + +## Troubleshooting + +### The Gitea install page appears instead of the login page + +The configuration file must exist at `/data/gitea/conf/app.ini` inside the container. If the mount path is wrong, Gitea starts with defaults and shows the installation wizard. Mount the file as shown in the Compose example and restart. + +### Push fails with an LFS or 403 error + +Confirm the credentials in `app.ini` match the RustFS credentials and that the `rustfs` hostname resolves inside the Compose network: + +```bash +docker compose logs gitea | grep -i minio +``` + +### Objects land in local storage instead of RustFS + +The `GITEA__storage__STORAGE_TYPE: minio` environment variable and the `[storage]` section of `app.ini` must agree. After changing either, restart Gitea and check the `Initialising` log lines again. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Gitea storage targets. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Gitea storage documentation](https://docs.gitea.com/administration/storage-configurations) to move packages, Actions artifacts, or individual storage sections to separate buckets. diff --git a/content/de/developer/integration/devops/images/rustfs-gitea-objects.png b/content/de/developer/integration/devops/images/rustfs-gitea-objects.png new file mode 100644 index 00000000..682f280c Binary files /dev/null and b/content/de/developer/integration/devops/images/rustfs-gitea-objects.png differ diff --git a/content/de/developer/integration/devops/index.md b/content/de/developer/integration/devops/index.md index 78391d8f..58fc491d 100644 --- a/content/de/developer/integration/devops/index.md +++ b/content/de/developer/integration/devops/index.md @@ -8,6 +8,7 @@ Nutzen Sie **RustFS** als Objektspeicher-Layer für DevOps-Plattformen und Infra ## Plattformen und Tools - [Elasticsearch](./elasticsearch.md) +- [Gitea](./gitea.md) - [Terraform](./terraform.md) Speichern Sie Artefakte, State und Telemetriedaten in dedizierten Buckets und beschränken Sie die Anmeldeinformationen auf die erforderlichen Bucket-Operationen. diff --git a/content/de/developer/integration/devops/meta.json b/content/de/developer/integration/devops/meta.json index a032acbd..61ad9fb8 100644 --- a/content/de/developer/integration/devops/meta.json +++ b/content/de/developer/integration/devops/meta.json @@ -2,6 +2,7 @@ "title": "DevOps", "pages": [ "elasticsearch", + "gitea", "terraform" ] } diff --git a/content/de/developer/integration/index.md b/content/de/developer/integration/index.md index 31beb5a1..868de8b2 100644 --- a/content/de/developer/integration/index.md +++ b/content/de/developer/integration/index.md @@ -8,11 +8,11 @@ Use this section to connect **RustFS** to infrastructure and application platfor ## Integration categories - [Reverse Proxy](./reverse-proxy/index.md) covers Nginx, Traefik, Caddy, and HAProxy. -- [Backup](./backup/index.md) covers Restic. +- [Backup](./backup/index.md) covers Restic and Longhorn. - [Datenanalyse](./big-data/index.md) covers Iceberg. - [Observability](./observability/index.md) covers OpenObserve. - [Others](./others/index.md) covers the community-driven capo SDK for Python. - [Registry](./registry/index.md) covers Harbor. -- [DevOps](./devops/index.md) covers Elasticsearch and Terraform. +- [DevOps](./devops/index.md) covers Elasticsearch, Gitea, and Terraform. Each guide identifies the RustFS endpoint and addressing requirements to use when configuring the integrating system. \ No newline at end of file diff --git a/content/de/developer/integration/observability/images/rustfs-thanos-blocks.png b/content/de/developer/integration/observability/images/rustfs-thanos-blocks.png new file mode 100644 index 00000000..188cb9cf Binary files /dev/null and b/content/de/developer/integration/observability/images/rustfs-thanos-blocks.png differ diff --git a/content/de/developer/integration/observability/index.md b/content/de/developer/integration/observability/index.md index 2290000a..3a1a4151 100644 --- a/content/de/developer/integration/observability/index.md +++ b/content/de/developer/integration/observability/index.md @@ -10,5 +10,6 @@ Nutzen Sie **RustFS** als Objektspeicher-Layer für Observability-Plattformen, d - [OpenObserve](./openobserve.md) - [Loki](./loki.md) - [Tempo](./tempo.md) +- [Thanos](./thanos.md) Speichern Sie Telemetriedaten in einem dedizierten Bucket und beschränken Sie die Anmeldeinformationen auf die erforderlichen Bucket-Operationen. diff --git a/content/de/developer/integration/observability/meta.json b/content/de/developer/integration/observability/meta.json index 7bbbe719..03f10c67 100644 --- a/content/de/developer/integration/observability/meta.json +++ b/content/de/developer/integration/observability/meta.json @@ -3,6 +3,7 @@ "pages": [ "openobserve", "loki", - "tempo" + "tempo", + "thanos" ] } diff --git a/content/de/developer/integration/observability/thanos.md b/content/de/developer/integration/observability/thanos.md new file mode 100644 index 00000000..856c31db --- /dev/null +++ b/content/de/developer/integration/observability/thanos.md @@ -0,0 +1,311 @@ +--- +title: "Thanos" +description: "Run Thanos with RustFS as the S3 object storage backend for Prometheus blocks, deployed with Docker Compose." +--- + +This guide connects [Thanos](https://github.com/thanos-io/thanos) — the highly available Prometheus setup with long-term storage — to **RustFS** as its object store. You will run Prometheus with a Thanos sidecar that uploads TSDB blocks to RustFS, then query the historical data back through a Store Gateway and a Query frontend. The workflow was verified with `thanosio/thanos:v0.37.2`, `prom/prometheus:v2.53.1`, and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with the Compose plugin. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Prom["Prometheus :9090"] -->|"blocks"| Sidecar["Thanos sidecar"] + Sidecar -->|"upload"| RustFS["RustFS :9000"] + Store["Store Gateway"] -->|"download"| RustFS + Query["Thanos Query"] -->|gRPC| Sidecar + Query -->|gRPC| Store +``` + +The sidecar watches the Prometheus TSDB directory and uploads every two-hour block to the `thanos-data` bucket in RustFS. The Store Gateway reads the same bucket and answers queries about historical blocks, so Query resolves both live data through the sidecar and old data through the Store Gateway. + +## 1. Create the project files + +Create a working directory: + +```bash +mkdir rustfs-thanos +cd rustfs-thanos +``` + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +RUSTFS_ACCESS_KEY= +RUSTFS_SECRET_KEY= +``` + +Use dedicated credentials for the `thanos-data` bucket. Do not commit `.env` to source control. + +Create the Prometheus configuration with an external label — Thanos requires it to deduplicate blocks: + +```yaml title="prometheus.yml" +global: + scrape_interval: 5s + external_labels: + monitor: rustfs-demo + +scrape_configs: + - job_name: prometheus + static_configs: + - targets: ["localhost:9090"] + - job_name: rustfs + metrics_path: /metrics + static_configs: + - targets: ["rustfs:9000"] +``` + +Create the Thanos object store configuration: + +```yaml title="bucket.yml" +type: S3 +config: + bucket: thanos-data + endpoint: rustfs:9000 + access_key: ${RUSTFS_ACCESS_KEY} + secret_key: ${RUSTFS_SECRET_KEY} + insecure: true +``` + +Thanos does not interpolate `.env` files itself. Before starting the stack, replace the placeholders with the same values you set in `.env`: + +```bash +sed -i.bak "s|\${RUSTFS_ACCESS_KEY}|$(grep RUSTFS_ACCESS_KEY .env | cut -d= -f2)|;s|\${RUSTFS_SECRET_KEY}|$(grep RUSTFS_SECRET_KEY .env | cut -d= -f2)|" bucket.yml +``` + +Create the Compose file: + +```yaml title="compose.yaml" +services: + rustfs: + image: rustfs/rustfs-x86-musl:v2.3.1 + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + RUSTFS_VOLUMES: /data + RUSTFS_ADDRESS: ":9000" + RUSTFS_CONSOLE_ADDRESS: ":9001" + RUSTFS_CONSOLE_ENABLE: "true" + volumes: + - rustfs-data:/data + ports: + - "9000:9000" + - "9001:9001" + healthcheck: + test: ["CMD", "curl", "-sf", "http://127.0.0.1:9000/health"] + interval: 10s + timeout: 5s + retries: 6 + start_period: 10s + networks: + - thanos + + create-bucket: + image: rustfs/rc:latest + depends_on: + rustfs: + condition: service_healthy + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + entrypoint: + - /bin/sh + - -c + - | + /usr/bin/rc alias set rustfs http://rustfs:9000 "$${RUSTFS_ACCESS_KEY}" "$${RUSTFS_SECRET_KEY}" + /usr/bin/rc mb --ignore-existing rustfs/thanos-data + networks: + - thanos + + prometheus: + image: prom/prometheus:v2.53.1 + command: + - --config.file=/etc/prometheus/prometheus.yml + - --storage.tsdb.path=/prometheus + - --storage.tsdb.min-block-duration=2h + - --storage.tsdb.max-block-duration=2h + - --web.enable-lifecycle + volumes: + - ./prometheus.yml:/etc/prometheus/prometheus.yml:ro + - prom-data:/prometheus + ports: + - "9090:9090" + networks: + - thanos + + sidecar: + image: thanosio/thanos:v0.37.2 + command: + - sidecar + - --tsdb.path=/prometheus + - --prometheus.url=http://prometheus:9090 + - --objstore.config-file=/etc/thanos/bucket.yml + volumes: + - ./bucket.yml:/etc/thanos/bucket.yml:ro + - prom-data:/prometheus + depends_on: + create-bucket: + condition: service_completed_successfully + networks: + - thanos + + store: + image: thanosio/thanos:v0.37.2 + command: + - store + - --objstore.config-file=/etc/thanos/bucket.yml + - --data-dir=/data + volumes: + - ./bucket.yml:/etc/thanos/bucket.yml:ro + - store-data:/data + depends_on: + create-bucket: + condition: service_completed_successfully + networks: + - thanos + + query: + image: thanosio/thanos:v0.37.2 + command: + - query + - --http-address=0.0.0.0:9090 + - --store=sidecar:10901 + - --store=store:10901 + ports: + - "9091:9090" + depends_on: + - sidecar + - store + networks: + - thanos + +networks: + thanos: + +volumes: + rustfs-data: + prom-data: + store-data: +``` + +The `--storage.tsdb.min-block-duration` and `--storage.tsdb.max-block-duration` flags disable Prometheus compaction. The sidecar refuses to ship blocks from a compacting TSDB because the local blocks would no longer match the uploaded ones. + +## 2. Validate and start the deployment + +Resolve the Compose file before starting containers: + +```bash +docker compose config +``` + +Start the stack and wait until the sidecar reports itself ready: + +```bash +docker compose up -d +docker compose logs sidecar | grep -m1 "status=ready" +``` + +Check that the sidecar picked up the Prometheus external labels: + +```bash +docker compose logs sidecar | grep "external labels" +``` + +The Thanos Query UI answers on `http://localhost:9091`, and the RustFS Console runs at `http://localhost:9001`. + +## 3. Upload a block to RustFS + +The sidecar uploads a block when Prometheus compacts one, which happens at a two-hour block boundary. To produce a block immediately, snapshot the TSDB through the admin API — with compaction disabled, the sidecar ships the head-block snapshot directly: + +```bash +curl -s -XPOST http://localhost:9090/api/v1/admin/tsdb/snapshot | head -c 200 +``` + +Wait for the upload, then check the shipper state inside Prometheus: + +```bash +sleep 60 +docker compose exec prometheus cat /prometheus/thanos.shipper.json +``` + +The `uploaded` list should contain a block ID: + +```json +{ + "version": 1, + "uploaded": [ + "01M31EPTZC5E0SETZTP0SPFY79" + ] +} +``` + +## 4. Query historical data from RustFS + +The Store Gateway periodically syncs the bucket. Confirm it downloaded the uploaded block: + +```bash +docker compose logs store | grep "loaded new block" +``` + +Query a series through the Query frontend over the block's time range: + +```bash +START=$(date -u -d '2 hours ago' +%s) +END=$(date -u +%s) +curl -s "http://localhost:9091/api/v1/query_range?query=up%7Bjob%3D%22prometheus%22%7D&start=$START&end=$END&step=30" | head -c 300 +``` + +The Store Gateway serves the response from the blocks it downloaded from RustFS, while the sidecar answers for the live head — both paths resolve through the same Query endpoint. + +## 5. Verify objects in RustFS + +List the bucket: + +```bash +docker compose exec rustfs /usr/bin/rc ls local/thanos-data/ -r +``` + +Each block is stored as three objects — the chunk files, the index, and `meta.json`: + +```text +01M31EPTZC5E0SETZTP0SPFY79/chunks/000001 +01M31EPTZC5E0SETZTP0SPFY79/index +01M31EPTZC5E0SETZTP0SPFY79/meta.json +``` + +![Thanos blocks stored in the RustFS Console](./images/rustfs-thanos-blocks.png) + +## 6. Stop or reset the deployment + +Stop the containers while keeping all data: + +```bash +docker compose down +``` + +The RustFS volume keeps the uploaded blocks, so the Store Gateway serves historical queries again after a restart. To delete everything, including the blocks in RustFS, add `--volumes`. + +## Troubleshooting + +### The sidecar logs `Compaction needs to be disabled` + +Prometheus must run with `--storage.tsdb.min-block-duration` equal to `--storage.tsdb.max-block-duration` — set both to `2h` as shown in the Compose file. Otherwise the sidecar cannot guarantee that local blocks stay unchanged and refuses to upload. + +### `The specified bucket does not exist` + +Thanos does not create buckets. Check that the `create-bucket` service completed successfully: + +```bash +docker compose logs create-bucket +``` + +### Queries return no historical data + +Confirm that the Store Gateway has loaded at least one block (`docker compose logs store | grep "loaded new block"`) and that your query time range falls inside the uploaded block's window — check the block's `meta.json` in the RustFS Console for `minTime` and `maxTime`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Thanos components. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Thanos documentation](https://thanos.io/tip/thanos/getting-started.md) to add Compactor, Ruler, or Receive for a production topology. diff --git a/content/en/developer/integration/backup/images/rustfs-longhorn-backups.png b/content/en/developer/integration/backup/images/rustfs-longhorn-backups.png new file mode 100644 index 00000000..899b1c8c Binary files /dev/null and b/content/en/developer/integration/backup/images/rustfs-longhorn-backups.png differ diff --git a/content/en/developer/integration/backup/index.md b/content/en/developer/integration/backup/index.md index b0c30303..9a5b5114 100644 --- a/content/en/developer/integration/backup/index.md +++ b/content/en/developer/integration/backup/index.md @@ -8,5 +8,6 @@ Use **RustFS** as the object storage backend for backup tools that store reposit ## Systems - [Restic](./restic.md) +- [Longhorn](./longhorn.md) Keep backup jobs in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations. \ No newline at end of file diff --git a/content/en/developer/integration/backup/longhorn.md b/content/en/developer/integration/backup/longhorn.md new file mode 100644 index 00000000..497f43cd --- /dev/null +++ b/content/en/developer/integration/backup/longhorn.md @@ -0,0 +1,293 @@ +--- +title: "Longhorn" +description: "Configure Longhorn to store Kubernetes volume backups in RustFS through its S3 backup target." +--- + +This guide connects [Longhorn](https://github.com/longhorn/longhorn) — the distributed block storage system for Kubernetes — to **RustFS** as its S3 backup target. You will configure the backup target, back up a volume that contains data, delete the volume, restore it from RustFS, and verify the data. The workflow was verified with Longhorn 1.9.0 on k3s (Kubernetes 1.30) and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need a Kubernetes cluster with Longhorn installed and `kubectl` access to it. This guide is intended for integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + App["Workload pod"] -->|"writes"| Vol["Longhorn volume"] + Vol -->|"snapshot"| Backup["Backup engine"] + Backup -->|"blocks + config"| RustFS["RustFS :9000"] +``` + +Longhorn stores backups in the `backupstore/volumes/` prefix of the target bucket as content-addressed blocks plus a small volume configuration file. Restores read those objects back into a new volume on any node that can reach RustFS. + +## 1. Create the backup bucket + +Create a dedicated bucket with the [`rc` client](https://github.com/rustfs/cli), using your own endpoint and credentials: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/longhorn-backups +``` + +Longhorn does not create the bucket, so this step must complete before the first backup. + +## 2. Configure the backup target + +Store the RustFS credentials in a secret in the `longhorn-system` namespace. `AWS_ENDPOINTS` must be an endpoint every node can reach — use the node IP or an internal load balancer address, not a port forward from your workstation: + +```bash +kubectl -n longhorn-system create secret generic rustfs-s3-secret \ + --from-literal=AWS_ACCESS_KEY_ID= \ + --from-literal=AWS_SECRET_ACCESS_KEY= \ + --from-literal=AWS_ENDPOINTS=http://:9000 +``` + +Longhorn 1.9 manages backup targets through the `BackupTarget` resource. Patch the `default` target with the RustFS bucket: + +```bash +kubectl -n longhorn-system patch backupTarget default --type merge -p ' +spec: + backupTargetURL: s3://longhorn-backups@us-east-1/ + credentialSecret: rustfs-s3-secret + pollInterval: 5m' +``` + +The `@us-east-1` segment is the region annotation in the S3 URL format; it does not need to match a real deployment region. + +Wait for the target to become available — it confirms that Longhorn reached the bucket through the secret: + +```bash +kubectl -n longhorn-system get backupTarget default +``` + +```text +NAME URL CREDENTIAL AVAILABLE LASTSYNCEDAT +default s3://longhorn-backups@us-east-1/ rustfs-s3-secret true 2026-09-21T08:36:55Z +``` + +## 3. Write data and create a backup + +Create a test volume with data in it. The following manifest creates a 1 GiB PVC and a pod that writes a marker file: + +```yaml title="demo.yaml" +apiVersion: v1 +kind: PersistentVolumeClaim +metadata: + name: demo-vol +spec: + accessModes: [ReadWriteOnce] + storageClassName: longhorn + resources: + requests: + storage: 1Gi +--- +apiVersion: v1 +kind: Pod +metadata: + name: demo-app +spec: + volumes: + - name: data + persistentVolumeClaim: + claimName: demo-vol + containers: + - name: app + image: busybox:1.36 + command: ["sh", "-c", "echo 'longhorn rustfs demo' > /data/hello.txt && sleep 3600"] + volumeMounts: + - name: data + mountPath: /data +``` + +Apply it and wait for the pod to run: + +```bash +kubectl apply -f demo.yaml +kubectl get pod demo-app +``` + +Create a snapshot and back it up. You can do this in the Longhorn UI (Volume → Snapshot → Backup) or declaratively: + +```bash +VOLUME=$(kubectl get pvc demo-vol -o jsonpath='{.spec.volumeName}') +kubectl -n longhorn-system apply -f - <" +EOF +``` + +The restore runs when the volume is first attached. Create a PVC bound to the restored volume through a static PV, then mount it: + +```bash +kubectl apply -f - < --type merge -p ' +spec: + disks: + default: + path: /var/lib/longhorn + allowScheduling: true' +``` + +### `failed to create backup ... missing input parameter` + +The backup ran before the default backup target was configured. Complete step 2, confirm `AVAILABLE` is `true`, and create the backup again. + +### The backup target never becomes available + +The secret must exist before the target syncs, and `AWS_ENDPOINTS` must be reachable from the nodes themselves. Check the `longhorn-manager` logs for S3 errors: + +```bash +kubectl -n longhorn-system logs -l app=longhorn-manager | grep -i s3 | tail +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional backup targets. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Longhorn backup documentation](https://longhorn.io/docs/1.9.0/backups-and-restore/) to configure recurring backup jobs and scheduled snapshots. diff --git a/content/en/developer/integration/backup/meta.json b/content/en/developer/integration/backup/meta.json index 9de69ffa..3ea20cbe 100644 --- a/content/en/developer/integration/backup/meta.json +++ b/content/en/developer/integration/backup/meta.json @@ -1,6 +1,7 @@ { "title": "Backup", "pages": [ - "restic" + "restic", + "longhorn" ] -} \ No newline at end of file +} diff --git a/content/en/developer/integration/big-data/images/rustfs-mlflow-artifacts.png b/content/en/developer/integration/big-data/images/rustfs-mlflow-artifacts.png new file mode 100644 index 00000000..545f5bf2 Binary files /dev/null and b/content/en/developer/integration/big-data/images/rustfs-mlflow-artifacts.png differ diff --git a/content/en/developer/integration/big-data/index.md b/content/en/developer/integration/big-data/index.md index 2325206d..f207d9b7 100644 --- a/content/en/developer/integration/big-data/index.md +++ b/content/en/developer/integration/big-data/index.md @@ -10,6 +10,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) - [Milvus](./milvus.md) +- [MLflow](./mlflow.md) - [DuckDB](./duckdb.md) - [InfluxDB](./influxdb.md) - [Spark](./spark.md) diff --git a/content/en/developer/integration/big-data/meta.json b/content/en/developer/integration/big-data/meta.json index 1d6f178f..87e29f91 100644 --- a/content/en/developer/integration/big-data/meta.json +++ b/content/en/developer/integration/big-data/meta.json @@ -4,6 +4,7 @@ "iceberg", "pyiceberg", "milvus", + "mlflow", "duckdb", "influxdb", "spark", diff --git a/content/en/developer/integration/big-data/mlflow.md b/content/en/developer/integration/big-data/mlflow.md new file mode 100644 index 00000000..46d3205e --- /dev/null +++ b/content/en/developer/integration/big-data/mlflow.md @@ -0,0 +1,244 @@ +--- +title: "MLflow" +description: "Run MLflow with RustFS as the S3 artifact store for experiment tracking, deployed with Docker Compose." +--- + +This guide connects [MLflow](https://github.com/mlflow/mlflow) — the experiment tracking and model registry platform — to **RustFS** as its S3 artifact store. You will start the MLflow tracking server with Docker Compose, log parameters, metrics, and artifacts from a training run, and verify that the artifacts are stored in RustFS. The workflow was verified with `ghcr.io/mlflow/mlflow:v2.22.1` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with the Compose plugin. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["Training script"] -->|"runs + metrics"| Server["MLflow server :5000"] + Client -->|"artifacts"| RustFS["RustFS :9000"] + Server -->|"metadata"| DB["SQLite"] +``` + +The tracking server keeps experiment and run metadata in SQLite and stores artifacts — model files, plots, reports — directly in RustFS through the `s3://` artifact root. The client needs the same RustFS credentials because it uploads artifacts itself, using boto3 and the `MLFLOW_S3_ENDPOINT_URL` setting. + +## 1. Create the project files + +Create a working directory: + +```bash +mkdir rustfs-mlflow +cd rustfs-mlflow +``` + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +RUSTFS_ACCESS_KEY= +RUSTFS_SECRET_KEY= +MLFLOW_BUCKET=my-bucket +``` + +Use dedicated credentials for the artifact bucket. Do not commit `.env` to source control. + +Create the Compose file: + +```yaml title="compose.yaml" +services: + rustfs: + image: rustfs/rustfs-x86-musl:v2.3.1 + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + RUSTFS_VOLUMES: /data + RUSTFS_ADDRESS: ":9000" + RUSTFS_CONSOLE_ADDRESS: ":9001" + RUSTFS_CONSOLE_ENABLE: "true" + volumes: + - rustfs-data:/data + ports: + - "9000:9000" + - "9001:9001" + healthcheck: + test: ["CMD", "curl", "-sf", "http://127.0.0.1:9000/health"] + interval: 10s + timeout: 5s + retries: 6 + start_period: 10s + networks: + - mlflow + + create-bucket: + image: rustfs/rc:latest + depends_on: + rustfs: + condition: service_healthy + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + MLFLOW_BUCKET: ${MLFLOW_BUCKET} + entrypoint: + - /bin/sh + - -c + - | + /usr/bin/rc alias set rustfs http://rustfs:9000 "$${RUSTFS_ACCESS_KEY}" "$${RUSTFS_SECRET_KEY}" + /usr/bin/rc mb --ignore-existing rustfs/$${MLFLOW_BUCKET} + networks: + - mlflow + + mlflow: + image: ghcr.io/mlflow/mlflow:v2.22.1 + command: + - server + - --backend-store-uri + - sqlite:////mlflow/mlflow.db + - --default-artifact-root + - s3://${MLFLOW_BUCKET}/mlflow-artifacts + - --host + - 0.0.0.0 + - --port + - "5000" + environment: + AWS_ACCESS_KEY_ID: ${RUSTFS_ACCESS_KEY} + AWS_SECRET_ACCESS_KEY: ${RUSTFS_SECRET_KEY} + MLFLOW_S3_ENDPOINT_URL: http://rustfs:9000 + AWS_DEFAULT_REGION: us-east-1 + depends_on: + create-bucket: + condition: service_completed_successfully + ports: + - "5000:5000" + volumes: + - mlflow-db:/mlflow + networks: + - mlflow + +networks: + mlflow: + +volumes: + rustfs-data: + mlflow-db: +``` + +The `create-bucket` service must run before the server starts because MLflow does not create the bucket. `MLFLOW_S3_ENDPOINT_URL` points the server's boto3 client at RustFS with path-style addressing. The SQLite database lives on a volume so experiment metadata survives restarts; production deployments should use a managed database backend instead. + +## 2. Validate and start the deployment + +Resolve the Compose file before starting containers: + +```bash +docker compose config +``` + +Start the services and wait for the tracking server: + +```bash +docker compose up -d +docker compose ps +``` + +The `create-bucket` service should exit with code `0`, and the MLflow UI should answer on `http://localhost:5000`: + +```bash +curl -sf http://localhost:5000/ >/dev/null && echo ready +``` + +Open the RustFS Console at `http://localhost:9001` to watch artifacts land in `my-bucket` during the next step. + +## 3. Log a training run + +Create the client script — it runs inside the MLflow image, which ships every dependency: + +```bash title="train_demo.py" {12} +import mlflow + +mlflow.set_tracking_uri("http://localhost:5000") +mlflow.set_experiment("rustfs-demo") + +with mlflow.start_run(run_name="rustfs-verify") as run: + mlflow.log_params({"model": "demo-regressor", "alpha": 0.5}) + for step in range(3): + mlflow.log_metric("rmse", 0.9 - step * 0.2, step=step) + with open("model-summary.txt", "w") as f: + f.write("demo model trained against RustFS artifact store\n") + mlflow.log_artifact("model-summary.txt", artifact_path="reports") + print("run_id:", run.info.run_id) + print("artifact_uri:", run.info.artifact_uri) +``` + +Copy the script into the running container and execute it with the service environment: + +```bash +docker compose cp train_demo.py mlflow:/tmp/train_demo.py +docker compose exec -w /tmp mlflow python train_demo.py +``` + +The `artifact_uri` prints as `s3://my-bucket/mlflow-artifacts///artifacts` — the artifact is uploaded straight to RustFS by the client. + +## 4. Verify artifacts in RustFS + +Read the artifact back through the tracking server, then list the same object in RustFS. Save the following as `verify.py`, replace `` with the identifier printed in step 3, and run it the same way: + +```bash title="verify.py" {5} +import mlflow + +mlflow.set_tracking_uri("http://localhost:5000") +client = mlflow.MlflowClient() +print([a.path for a in client.list_artifacts("", "reports")]) +path = client.download_artifacts("", "reports/model-summary.txt") +print(open(path).read()) +``` + +The download reads the object from RustFS through the tracking server. Then confirm the objects in the bucket: + +```bash +docker compose exec rustfs /usr/bin/rc ls local/my-bucket/mlflow-artifacts/ -r +``` + +The output should include the artifact object: + +```text +mlflow-artifacts/1/a1aece9243504f2680a52fba0c32765f/artifacts/reports/model-summary.txt +``` + +![MLflow artifacts stored in the RustFS Console](./images/rustfs-mlflow-artifacts.png) + +Runs, parameters, and metrics survive a server restart because they are stored in SQLite, while the artifacts stay in RustFS: + +```bash +docker compose restart mlflow +docker compose exec -w /tmp mlflow python verify.py +``` + +The `list_artifacts` call succeeds again against the restarted server, reading from the same objects in RustFS. + +## 5. Stop or reset the deployment + +Stop the containers while keeping all data: + +```bash +docker compose down +``` + +The RustFS volume keeps the artifact objects and the MLflow volume keeps the metadata database. To delete everything, including the artifacts in RustFS, add `--volumes`. + +## Troubleshooting + +### `ModuleNotFoundError: No module named 'boto3'` when logging artifacts + +The client performing `log_artifact` needs boto3 because it uploads directly to S3. Install it in the environment that runs the training script, or run the script inside the MLflow image as shown in step 3. + +### `AccessDenied` or connection errors during artifact upload + +Confirm that `MLFLOW_S3_ENDPOINT_URL` is set in the client environment — without it, boto3 sends requests to real AWS S3. The endpoint must be reachable from the machine running the training script; use `http://localhost:9000` outside the Compose network and `http://rustfs:9000` inside it. + +### The server fails to start with a bucket error + +The artifact bucket must exist before the server starts. Check the `create-bucket` service logs: + +```bash +docker compose logs create-bucket +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional MLflow operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [MLflow documentation](https://mlflow.org/docs/latest/) to add a model registry or move the metadata store to a managed database. diff --git a/content/en/developer/integration/devops/gitea.md b/content/en/developer/integration/devops/gitea.md new file mode 100644 index 00000000..239b6ad2 --- /dev/null +++ b/content/en/developer/integration/devops/gitea.md @@ -0,0 +1,236 @@ +--- +title: "Gitea" +description: "Run Gitea with RustFS as the S3 storage backend for LFS objects and attachments, deployed with Docker Compose." +--- + +This guide connects [Gitea](https://github.com/go-gitea/gitea) — the self-hosted Git service — to **RustFS** through Gitea's `minio` storage type. You will start Gitea with Docker Compose, create a repository, push a Git LFS object, attach a file to an issue, and verify that both the LFS object and the attachment are stored in RustFS. The workflow was verified with `gitea/gitea:1.24.4` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with the Compose plugin and the `git` and `git-lfs` clients on your workstation. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Dev["Git + LFS client"] -->|"git push / git-lfs"| Gitea["Gitea :3000"] + Gitea -->|"LFS + attachments"| RustFS["RustFS :9000"] +``` + +Gitea stores the Git repository itself on its local disk, while the `minio` storage type routes large files — LFS objects, issue attachments, avatars, repository archives, packages, and Actions artifacts — to the `gitea-data` bucket in RustFS. Gitea creates the bucket on startup if it does not exist. + +## 1. Create the project files + +Create a working directory: + +```bash +mkdir rustfs-gitea +cd rustfs-gitea +``` + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +RUSTFS_ACCESS_KEY= +RUSTFS_SECRET_KEY= +``` + +Use dedicated credentials for the `gitea-data` bucket. Do not commit `.env` to source control. + +Create the Gitea configuration and replace both credential placeholders with the same values you set in `.env` — the global `minio` storage type applies to LFS, attachments, avatars, repository archives, packages, and Actions artifacts: + +```ini title="app.ini" +APP_NAME = RustFS Gitea +RUN_MODE = prod +WORK_PATH = /data/gitea + +[server] +DOMAIN = localhost +ROOT_URL = http://localhost:3000/ +HTTP_PORT = 3000 +LFS_START_SERVER = true + +[database] +DB_TYPE = sqlite3 +PATH = /data/gitea/gitea.db + +[storage] +STORAGE_TYPE = minio +MINIO_ENDPOINT = rustfs:9000 +MINIO_ACCESS_KEY_ID = +MINIO_SECRET_ACCESS_KEY = +MINIO_BUCKET = gitea-data +MINIO_LOCATION = us-east-1 +MINIO_USE_SSL = false + +[log] +MODE = console +LEVEL = info + +[security] +INSTALL_LOCK = true +SECRET_KEY = change-me-to-a-random-string +``` + +`LFS_START_SERVER` enables the Git LFS HTTP API. `MINIO_ENDPOINT` uses the Compose-internal hostname `rustfs`; the endpoint is reached with path-style requests by default. + +Create the Compose file: + +```yaml title="compose.yaml" +services: + rustfs: + image: rustfs/rustfs-x86-musl:v2.3.1 + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + RUSTFS_VOLUMES: /data + RUSTFS_ADDRESS: ":9000" + RUSTFS_CONSOLE_ADDRESS: ":9001" + RUSTFS_CONSOLE_ENABLE: "true" + volumes: + - rustfs-data:/data + ports: + - "9000:9000" + - "9001:9001" + healthcheck: + test: ["CMD", "curl", "-sf", "http://127.0.0.1:9000/health"] + interval: 10s + timeout: 5s + retries: 6 + start_period: 10s + networks: + - gitea + + gitea: + image: gitea/gitea:1.24.4 + depends_on: + rustfs: + condition: service_healthy + environment: + USER_UID: "1000" + USER_GID: "1000" + volumes: + - gitea-data:/data + - ./app.ini:/data/gitea/conf/app.ini:ro + ports: + - "3000:3000" + networks: + - gitea + +networks: + gitea: + +volumes: + rustfs-data: + gitea-data: +``` + +The volume-backed `/data` directory keeps the SQLite database and the Git repositories across container restarts, while LFS objects and attachments live in RustFS. + +## 2. Validate and start the deployment + +Resolve the Compose file before starting containers: + +```bash +docker compose config +``` + +Start the services: + +```bash +docker compose up -d +docker compose ps +``` + +Watch the Gitea log until every storage backend reports the Minio type: + +```bash +docker compose logs gitea | grep "Initialising" +``` + +The output should list `Attachment`, `Avatar`, `LFS`, and the remaining storage sections, each followed by a `Creating Minio storage at rustfs:9000:gitea-data` line. + +Open `http://localhost:3000` and create the administrator account, then open the RustFS Console at `http://localhost:9001` — the `gitea-data` bucket appears after the first storage operation. + +## 3. Push a Git LFS object + +Create a repository named `rustfs-demo` in the Gitea web UI, then push an LFS-tracked file from your workstation: + +```bash +mkdir lfs-demo && cd lfs-demo +git init +git config user.email you@example.com +git config user.name you +git lfs install +git lfs track "*.bin" +git add .gitattributes +dd if=/dev/urandom of=dataset.bin bs=1M count=8 +git add dataset.bin +git commit -m "add LFS dataset" +git remote add origin http://localhost:3000//rustfs-demo.git +git push origin main +``` + +`git push` uploads the LFS object through the Gitea LFS API, which writes it to RustFS. Clone the repository into a second directory and run `git lfs pull` — the downloaded `dataset.bin` must be byte-identical to the original: + +```bash +sha256sum dataset.bin +cd ../lfs-demo-clone && git lfs pull && sha256sum dataset.bin +``` + +Both checksums match because both clients read the object from RustFS. + +## 4. Attach a file to an issue + +Open the `rustfs-demo` repository, create an issue, and attach a small text file through the issue form. Gitea stores the upload as `attachments//` in the `gitea-data` bucket and serves downloads through `/attachments/`. + +## 5. Verify objects in RustFS + +List the bucket with the [`rc` client](https://github.com/rustfs/cli): + +```bash +docker compose exec rustfs /usr/bin/rc ls local/gitea-data/ -r +``` + +The output should include the LFS object under `lfs/` and the attachment under `attachments/`: + +```text +attachments/9/2/92fdd48d-531c-4cba-8b3f-4e2004a10fc7 +lfs/37/76/6ddfc07e803de58a69328db9a58a07cf7080ddde55c155a7531bc650a000 +``` + +The LFS object key is the SHA-256 content hash used by the Git LFS protocol. + +![Gitea LFS and attachment objects in the RustFS Console](./images/rustfs-gitea-objects.png) + +## 6. Stop or reset the deployment + +Stop the containers while keeping all data: + +```bash +docker compose down +``` + +The RustFS volume keeps the `gitea-data` bucket, so LFS objects and attachments survive a restart. To delete everything, including the objects in RustFS, add `--volumes`. + +## Troubleshooting + +### The Gitea install page appears instead of the login page + +The configuration file must exist at `/data/gitea/conf/app.ini` inside the container. If the mount path is wrong, Gitea starts with defaults and shows the installation wizard. Mount the file as shown in the Compose example and restart. + +### Push fails with an LFS or 403 error + +Confirm the credentials in `app.ini` match the RustFS credentials and that the `rustfs` hostname resolves inside the Compose network: + +```bash +docker compose logs gitea | grep -i minio +``` + +### Objects land in local storage instead of RustFS + +The `GITEA__storage__STORAGE_TYPE: minio` environment variable and the `[storage]` section of `app.ini` must agree. After changing either, restart Gitea and check the `Initialising` log lines again. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Gitea storage targets. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Gitea storage documentation](https://docs.gitea.com/administration/storage-configurations) to move packages, Actions artifacts, or individual storage sections to separate buckets. diff --git a/content/en/developer/integration/devops/images/rustfs-gitea-objects.png b/content/en/developer/integration/devops/images/rustfs-gitea-objects.png new file mode 100644 index 00000000..682f280c Binary files /dev/null and b/content/en/developer/integration/devops/images/rustfs-gitea-objects.png differ diff --git a/content/en/developer/integration/devops/index.md b/content/en/developer/integration/devops/index.md index a4a85574..95f6261e 100644 --- a/content/en/developer/integration/devops/index.md +++ b/content/en/developer/integration/devops/index.md @@ -8,6 +8,7 @@ Use **RustFS** as the object storage layer for DevOps platforms and infrastructu ## Platforms - [Elasticsearch](./elasticsearch.md) +- [Gitea](./gitea.md) - [Terraform](./terraform.md) Keep artifacts, state, and telemetry in dedicated buckets, and use credentials scoped to the required bucket operations. diff --git a/content/en/developer/integration/devops/meta.json b/content/en/developer/integration/devops/meta.json index a032acbd..61ad9fb8 100644 --- a/content/en/developer/integration/devops/meta.json +++ b/content/en/developer/integration/devops/meta.json @@ -2,6 +2,7 @@ "title": "DevOps", "pages": [ "elasticsearch", + "gitea", "terraform" ] } diff --git a/content/en/developer/integration/index.md b/content/en/developer/integration/index.md index dadfe77b..5e604df9 100644 --- a/content/en/developer/integration/index.md +++ b/content/en/developer/integration/index.md @@ -8,11 +8,11 @@ Use this section to connect **RustFS** to infrastructure and application platfor ## Integration categories - [Reverse Proxy](./reverse-proxy/index.md) covers Nginx, Traefik, Caddy, and HAProxy. -- [Backup](./backup/index.md) covers Restic. +- [Backup](./backup/index.md) covers Restic and Longhorn. - [Data Analytics](./big-data/index.md) covers Iceberg. - [Observability](./observability/index.md) covers OpenObserve. - [Others](./others/index.md) covers the community-driven capo SDK for Python. - [Registry](./registry/index.md) covers Harbor. -- [DevOps](./devops/index.md) covers Elasticsearch and Terraform. +- [DevOps](./devops/index.md) covers Elasticsearch, Gitea, and Terraform. Each guide identifies the RustFS endpoint and addressing requirements to use when configuring the integrating system. \ No newline at end of file diff --git a/content/en/developer/integration/observability/images/rustfs-thanos-blocks.png b/content/en/developer/integration/observability/images/rustfs-thanos-blocks.png new file mode 100644 index 00000000..188cb9cf Binary files /dev/null and b/content/en/developer/integration/observability/images/rustfs-thanos-blocks.png differ diff --git a/content/en/developer/integration/observability/index.md b/content/en/developer/integration/observability/index.md index c91dc8b9..dbdbea68 100644 --- a/content/en/developer/integration/observability/index.md +++ b/content/en/developer/integration/observability/index.md @@ -10,5 +10,6 @@ Use **RustFS** as the object storage layer for observability platforms that supp - [OpenObserve](./openobserve.md) - [Loki](./loki.md) - [Tempo](./tempo.md) +- [Thanos](./thanos.md) Keep telemetry data in a dedicated bucket, and use credentials scoped to the required bucket operations. diff --git a/content/en/developer/integration/observability/meta.json b/content/en/developer/integration/observability/meta.json index 7bbbe719..03f10c67 100644 --- a/content/en/developer/integration/observability/meta.json +++ b/content/en/developer/integration/observability/meta.json @@ -3,6 +3,7 @@ "pages": [ "openobserve", "loki", - "tempo" + "tempo", + "thanos" ] } diff --git a/content/en/developer/integration/observability/thanos.md b/content/en/developer/integration/observability/thanos.md new file mode 100644 index 00000000..856c31db --- /dev/null +++ b/content/en/developer/integration/observability/thanos.md @@ -0,0 +1,311 @@ +--- +title: "Thanos" +description: "Run Thanos with RustFS as the S3 object storage backend for Prometheus blocks, deployed with Docker Compose." +--- + +This guide connects [Thanos](https://github.com/thanos-io/thanos) — the highly available Prometheus setup with long-term storage — to **RustFS** as its object store. You will run Prometheus with a Thanos sidecar that uploads TSDB blocks to RustFS, then query the historical data back through a Store Gateway and a Query frontend. The workflow was verified with `thanosio/thanos:v0.37.2`, `prom/prometheus:v2.53.1`, and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with the Compose plugin. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Prom["Prometheus :9090"] -->|"blocks"| Sidecar["Thanos sidecar"] + Sidecar -->|"upload"| RustFS["RustFS :9000"] + Store["Store Gateway"] -->|"download"| RustFS + Query["Thanos Query"] -->|gRPC| Sidecar + Query -->|gRPC| Store +``` + +The sidecar watches the Prometheus TSDB directory and uploads every two-hour block to the `thanos-data` bucket in RustFS. The Store Gateway reads the same bucket and answers queries about historical blocks, so Query resolves both live data through the sidecar and old data through the Store Gateway. + +## 1. Create the project files + +Create a working directory: + +```bash +mkdir rustfs-thanos +cd rustfs-thanos +``` + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +RUSTFS_ACCESS_KEY= +RUSTFS_SECRET_KEY= +``` + +Use dedicated credentials for the `thanos-data` bucket. Do not commit `.env` to source control. + +Create the Prometheus configuration with an external label — Thanos requires it to deduplicate blocks: + +```yaml title="prometheus.yml" +global: + scrape_interval: 5s + external_labels: + monitor: rustfs-demo + +scrape_configs: + - job_name: prometheus + static_configs: + - targets: ["localhost:9090"] + - job_name: rustfs + metrics_path: /metrics + static_configs: + - targets: ["rustfs:9000"] +``` + +Create the Thanos object store configuration: + +```yaml title="bucket.yml" +type: S3 +config: + bucket: thanos-data + endpoint: rustfs:9000 + access_key: ${RUSTFS_ACCESS_KEY} + secret_key: ${RUSTFS_SECRET_KEY} + insecure: true +``` + +Thanos does not interpolate `.env` files itself. Before starting the stack, replace the placeholders with the same values you set in `.env`: + +```bash +sed -i.bak "s|\${RUSTFS_ACCESS_KEY}|$(grep RUSTFS_ACCESS_KEY .env | cut -d= -f2)|;s|\${RUSTFS_SECRET_KEY}|$(grep RUSTFS_SECRET_KEY .env | cut -d= -f2)|" bucket.yml +``` + +Create the Compose file: + +```yaml title="compose.yaml" +services: + rustfs: + image: rustfs/rustfs-x86-musl:v2.3.1 + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + RUSTFS_VOLUMES: /data + RUSTFS_ADDRESS: ":9000" + RUSTFS_CONSOLE_ADDRESS: ":9001" + RUSTFS_CONSOLE_ENABLE: "true" + volumes: + - rustfs-data:/data + ports: + - "9000:9000" + - "9001:9001" + healthcheck: + test: ["CMD", "curl", "-sf", "http://127.0.0.1:9000/health"] + interval: 10s + timeout: 5s + retries: 6 + start_period: 10s + networks: + - thanos + + create-bucket: + image: rustfs/rc:latest + depends_on: + rustfs: + condition: service_healthy + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + entrypoint: + - /bin/sh + - -c + - | + /usr/bin/rc alias set rustfs http://rustfs:9000 "$${RUSTFS_ACCESS_KEY}" "$${RUSTFS_SECRET_KEY}" + /usr/bin/rc mb --ignore-existing rustfs/thanos-data + networks: + - thanos + + prometheus: + image: prom/prometheus:v2.53.1 + command: + - --config.file=/etc/prometheus/prometheus.yml + - --storage.tsdb.path=/prometheus + - --storage.tsdb.min-block-duration=2h + - --storage.tsdb.max-block-duration=2h + - --web.enable-lifecycle + volumes: + - ./prometheus.yml:/etc/prometheus/prometheus.yml:ro + - prom-data:/prometheus + ports: + - "9090:9090" + networks: + - thanos + + sidecar: + image: thanosio/thanos:v0.37.2 + command: + - sidecar + - --tsdb.path=/prometheus + - --prometheus.url=http://prometheus:9090 + - --objstore.config-file=/etc/thanos/bucket.yml + volumes: + - ./bucket.yml:/etc/thanos/bucket.yml:ro + - prom-data:/prometheus + depends_on: + create-bucket: + condition: service_completed_successfully + networks: + - thanos + + store: + image: thanosio/thanos:v0.37.2 + command: + - store + - --objstore.config-file=/etc/thanos/bucket.yml + - --data-dir=/data + volumes: + - ./bucket.yml:/etc/thanos/bucket.yml:ro + - store-data:/data + depends_on: + create-bucket: + condition: service_completed_successfully + networks: + - thanos + + query: + image: thanosio/thanos:v0.37.2 + command: + - query + - --http-address=0.0.0.0:9090 + - --store=sidecar:10901 + - --store=store:10901 + ports: + - "9091:9090" + depends_on: + - sidecar + - store + networks: + - thanos + +networks: + thanos: + +volumes: + rustfs-data: + prom-data: + store-data: +``` + +The `--storage.tsdb.min-block-duration` and `--storage.tsdb.max-block-duration` flags disable Prometheus compaction. The sidecar refuses to ship blocks from a compacting TSDB because the local blocks would no longer match the uploaded ones. + +## 2. Validate and start the deployment + +Resolve the Compose file before starting containers: + +```bash +docker compose config +``` + +Start the stack and wait until the sidecar reports itself ready: + +```bash +docker compose up -d +docker compose logs sidecar | grep -m1 "status=ready" +``` + +Check that the sidecar picked up the Prometheus external labels: + +```bash +docker compose logs sidecar | grep "external labels" +``` + +The Thanos Query UI answers on `http://localhost:9091`, and the RustFS Console runs at `http://localhost:9001`. + +## 3. Upload a block to RustFS + +The sidecar uploads a block when Prometheus compacts one, which happens at a two-hour block boundary. To produce a block immediately, snapshot the TSDB through the admin API — with compaction disabled, the sidecar ships the head-block snapshot directly: + +```bash +curl -s -XPOST http://localhost:9090/api/v1/admin/tsdb/snapshot | head -c 200 +``` + +Wait for the upload, then check the shipper state inside Prometheus: + +```bash +sleep 60 +docker compose exec prometheus cat /prometheus/thanos.shipper.json +``` + +The `uploaded` list should contain a block ID: + +```json +{ + "version": 1, + "uploaded": [ + "01M31EPTZC5E0SETZTP0SPFY79" + ] +} +``` + +## 4. Query historical data from RustFS + +The Store Gateway periodically syncs the bucket. Confirm it downloaded the uploaded block: + +```bash +docker compose logs store | grep "loaded new block" +``` + +Query a series through the Query frontend over the block's time range: + +```bash +START=$(date -u -d '2 hours ago' +%s) +END=$(date -u +%s) +curl -s "http://localhost:9091/api/v1/query_range?query=up%7Bjob%3D%22prometheus%22%7D&start=$START&end=$END&step=30" | head -c 300 +``` + +The Store Gateway serves the response from the blocks it downloaded from RustFS, while the sidecar answers for the live head — both paths resolve through the same Query endpoint. + +## 5. Verify objects in RustFS + +List the bucket: + +```bash +docker compose exec rustfs /usr/bin/rc ls local/thanos-data/ -r +``` + +Each block is stored as three objects — the chunk files, the index, and `meta.json`: + +```text +01M31EPTZC5E0SETZTP0SPFY79/chunks/000001 +01M31EPTZC5E0SETZTP0SPFY79/index +01M31EPTZC5E0SETZTP0SPFY79/meta.json +``` + +![Thanos blocks stored in the RustFS Console](./images/rustfs-thanos-blocks.png) + +## 6. Stop or reset the deployment + +Stop the containers while keeping all data: + +```bash +docker compose down +``` + +The RustFS volume keeps the uploaded blocks, so the Store Gateway serves historical queries again after a restart. To delete everything, including the blocks in RustFS, add `--volumes`. + +## Troubleshooting + +### The sidecar logs `Compaction needs to be disabled` + +Prometheus must run with `--storage.tsdb.min-block-duration` equal to `--storage.tsdb.max-block-duration` — set both to `2h` as shown in the Compose file. Otherwise the sidecar cannot guarantee that local blocks stay unchanged and refuses to upload. + +### `The specified bucket does not exist` + +Thanos does not create buckets. Check that the `create-bucket` service completed successfully: + +```bash +docker compose logs create-bucket +``` + +### Queries return no historical data + +Confirm that the Store Gateway has loaded at least one block (`docker compose logs store | grep "loaded new block"`) and that your query time range falls inside the uploaded block's window — check the block's `meta.json` in the RustFS Console for `minTime` and `maxTime`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Thanos components. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Thanos documentation](https://thanos.io/tip/thanos/getting-started.md) to add Compactor, Ruler, or Receive for a production topology. diff --git a/content/fr/developer/integration/backup/images/rustfs-longhorn-backups.png b/content/fr/developer/integration/backup/images/rustfs-longhorn-backups.png new file mode 100644 index 00000000..899b1c8c Binary files /dev/null and b/content/fr/developer/integration/backup/images/rustfs-longhorn-backups.png differ diff --git a/content/fr/developer/integration/backup/index.md b/content/fr/developer/integration/backup/index.md index 19e82d4c..b29cbf60 100644 --- a/content/fr/developer/integration/backup/index.md +++ b/content/fr/developer/integration/backup/index.md @@ -8,5 +8,6 @@ Utilisez **RustFS** comme backend de stockage objet pour les outils de sauvegard ## Systèmes - [Restic](./restic.md) +- [Longhorn](./longhorn.md) Conservez les tâches de sauvegarde dans un compartiment et un préfixe dédiés, et utilisez des identifiants limités aux opérations de compartiment nécessaires. \ No newline at end of file diff --git a/content/fr/developer/integration/backup/longhorn.md b/content/fr/developer/integration/backup/longhorn.md new file mode 100644 index 00000000..497f43cd --- /dev/null +++ b/content/fr/developer/integration/backup/longhorn.md @@ -0,0 +1,293 @@ +--- +title: "Longhorn" +description: "Configure Longhorn to store Kubernetes volume backups in RustFS through its S3 backup target." +--- + +This guide connects [Longhorn](https://github.com/longhorn/longhorn) — the distributed block storage system for Kubernetes — to **RustFS** as its S3 backup target. You will configure the backup target, back up a volume that contains data, delete the volume, restore it from RustFS, and verify the data. The workflow was verified with Longhorn 1.9.0 on k3s (Kubernetes 1.30) and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need a Kubernetes cluster with Longhorn installed and `kubectl` access to it. This guide is intended for integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + App["Workload pod"] -->|"writes"| Vol["Longhorn volume"] + Vol -->|"snapshot"| Backup["Backup engine"] + Backup -->|"blocks + config"| RustFS["RustFS :9000"] +``` + +Longhorn stores backups in the `backupstore/volumes/` prefix of the target bucket as content-addressed blocks plus a small volume configuration file. Restores read those objects back into a new volume on any node that can reach RustFS. + +## 1. Create the backup bucket + +Create a dedicated bucket with the [`rc` client](https://github.com/rustfs/cli), using your own endpoint and credentials: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/longhorn-backups +``` + +Longhorn does not create the bucket, so this step must complete before the first backup. + +## 2. Configure the backup target + +Store the RustFS credentials in a secret in the `longhorn-system` namespace. `AWS_ENDPOINTS` must be an endpoint every node can reach — use the node IP or an internal load balancer address, not a port forward from your workstation: + +```bash +kubectl -n longhorn-system create secret generic rustfs-s3-secret \ + --from-literal=AWS_ACCESS_KEY_ID= \ + --from-literal=AWS_SECRET_ACCESS_KEY= \ + --from-literal=AWS_ENDPOINTS=http://:9000 +``` + +Longhorn 1.9 manages backup targets through the `BackupTarget` resource. Patch the `default` target with the RustFS bucket: + +```bash +kubectl -n longhorn-system patch backupTarget default --type merge -p ' +spec: + backupTargetURL: s3://longhorn-backups@us-east-1/ + credentialSecret: rustfs-s3-secret + pollInterval: 5m' +``` + +The `@us-east-1` segment is the region annotation in the S3 URL format; it does not need to match a real deployment region. + +Wait for the target to become available — it confirms that Longhorn reached the bucket through the secret: + +```bash +kubectl -n longhorn-system get backupTarget default +``` + +```text +NAME URL CREDENTIAL AVAILABLE LASTSYNCEDAT +default s3://longhorn-backups@us-east-1/ rustfs-s3-secret true 2026-09-21T08:36:55Z +``` + +## 3. Write data and create a backup + +Create a test volume with data in it. The following manifest creates a 1 GiB PVC and a pod that writes a marker file: + +```yaml title="demo.yaml" +apiVersion: v1 +kind: PersistentVolumeClaim +metadata: + name: demo-vol +spec: + accessModes: [ReadWriteOnce] + storageClassName: longhorn + resources: + requests: + storage: 1Gi +--- +apiVersion: v1 +kind: Pod +metadata: + name: demo-app +spec: + volumes: + - name: data + persistentVolumeClaim: + claimName: demo-vol + containers: + - name: app + image: busybox:1.36 + command: ["sh", "-c", "echo 'longhorn rustfs demo' > /data/hello.txt && sleep 3600"] + volumeMounts: + - name: data + mountPath: /data +``` + +Apply it and wait for the pod to run: + +```bash +kubectl apply -f demo.yaml +kubectl get pod demo-app +``` + +Create a snapshot and back it up. You can do this in the Longhorn UI (Volume → Snapshot → Backup) or declaratively: + +```bash +VOLUME=$(kubectl get pvc demo-vol -o jsonpath='{.spec.volumeName}') +kubectl -n longhorn-system apply -f - <" +EOF +``` + +The restore runs when the volume is first attached. Create a PVC bound to the restored volume through a static PV, then mount it: + +```bash +kubectl apply -f - < --type merge -p ' +spec: + disks: + default: + path: /var/lib/longhorn + allowScheduling: true' +``` + +### `failed to create backup ... missing input parameter` + +The backup ran before the default backup target was configured. Complete step 2, confirm `AVAILABLE` is `true`, and create the backup again. + +### The backup target never becomes available + +The secret must exist before the target syncs, and `AWS_ENDPOINTS` must be reachable from the nodes themselves. Check the `longhorn-manager` logs for S3 errors: + +```bash +kubectl -n longhorn-system logs -l app=longhorn-manager | grep -i s3 | tail +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional backup targets. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Longhorn backup documentation](https://longhorn.io/docs/1.9.0/backups-and-restore/) to configure recurring backup jobs and scheduled snapshots. diff --git a/content/fr/developer/integration/backup/meta.json b/content/fr/developer/integration/backup/meta.json index 01ed058d..b9923706 100644 --- a/content/fr/developer/integration/backup/meta.json +++ b/content/fr/developer/integration/backup/meta.json @@ -1,6 +1,7 @@ { "title": "Sauvegarde", "pages": [ - "restic" + "restic", + "longhorn" ] -} \ No newline at end of file +} diff --git a/content/fr/developer/integration/big-data/images/rustfs-mlflow-artifacts.png b/content/fr/developer/integration/big-data/images/rustfs-mlflow-artifacts.png new file mode 100644 index 00000000..545f5bf2 Binary files /dev/null and b/content/fr/developer/integration/big-data/images/rustfs-mlflow-artifacts.png differ diff --git a/content/fr/developer/integration/big-data/index.md b/content/fr/developer/integration/big-data/index.md index 11b4660d..e0d11e44 100644 --- a/content/fr/developer/integration/big-data/index.md +++ b/content/fr/developer/integration/big-data/index.md @@ -10,6 +10,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) - [Milvus](./milvus.md) +- [MLflow](./mlflow.md) - [DuckDB](./duckdb.md) - [InfluxDB](./influxdb.md) - [Spark](./spark.md) diff --git a/content/fr/developer/integration/big-data/meta.json b/content/fr/developer/integration/big-data/meta.json index 1d6f178f..87e29f91 100644 --- a/content/fr/developer/integration/big-data/meta.json +++ b/content/fr/developer/integration/big-data/meta.json @@ -4,6 +4,7 @@ "iceberg", "pyiceberg", "milvus", + "mlflow", "duckdb", "influxdb", "spark", diff --git a/content/fr/developer/integration/big-data/mlflow.md b/content/fr/developer/integration/big-data/mlflow.md new file mode 100644 index 00000000..46d3205e --- /dev/null +++ b/content/fr/developer/integration/big-data/mlflow.md @@ -0,0 +1,244 @@ +--- +title: "MLflow" +description: "Run MLflow with RustFS as the S3 artifact store for experiment tracking, deployed with Docker Compose." +--- + +This guide connects [MLflow](https://github.com/mlflow/mlflow) — the experiment tracking and model registry platform — to **RustFS** as its S3 artifact store. You will start the MLflow tracking server with Docker Compose, log parameters, metrics, and artifacts from a training run, and verify that the artifacts are stored in RustFS. The workflow was verified with `ghcr.io/mlflow/mlflow:v2.22.1` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with the Compose plugin. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["Training script"] -->|"runs + metrics"| Server["MLflow server :5000"] + Client -->|"artifacts"| RustFS["RustFS :9000"] + Server -->|"metadata"| DB["SQLite"] +``` + +The tracking server keeps experiment and run metadata in SQLite and stores artifacts — model files, plots, reports — directly in RustFS through the `s3://` artifact root. The client needs the same RustFS credentials because it uploads artifacts itself, using boto3 and the `MLFLOW_S3_ENDPOINT_URL` setting. + +## 1. Create the project files + +Create a working directory: + +```bash +mkdir rustfs-mlflow +cd rustfs-mlflow +``` + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +RUSTFS_ACCESS_KEY= +RUSTFS_SECRET_KEY= +MLFLOW_BUCKET=my-bucket +``` + +Use dedicated credentials for the artifact bucket. Do not commit `.env` to source control. + +Create the Compose file: + +```yaml title="compose.yaml" +services: + rustfs: + image: rustfs/rustfs-x86-musl:v2.3.1 + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + RUSTFS_VOLUMES: /data + RUSTFS_ADDRESS: ":9000" + RUSTFS_CONSOLE_ADDRESS: ":9001" + RUSTFS_CONSOLE_ENABLE: "true" + volumes: + - rustfs-data:/data + ports: + - "9000:9000" + - "9001:9001" + healthcheck: + test: ["CMD", "curl", "-sf", "http://127.0.0.1:9000/health"] + interval: 10s + timeout: 5s + retries: 6 + start_period: 10s + networks: + - mlflow + + create-bucket: + image: rustfs/rc:latest + depends_on: + rustfs: + condition: service_healthy + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + MLFLOW_BUCKET: ${MLFLOW_BUCKET} + entrypoint: + - /bin/sh + - -c + - | + /usr/bin/rc alias set rustfs http://rustfs:9000 "$${RUSTFS_ACCESS_KEY}" "$${RUSTFS_SECRET_KEY}" + /usr/bin/rc mb --ignore-existing rustfs/$${MLFLOW_BUCKET} + networks: + - mlflow + + mlflow: + image: ghcr.io/mlflow/mlflow:v2.22.1 + command: + - server + - --backend-store-uri + - sqlite:////mlflow/mlflow.db + - --default-artifact-root + - s3://${MLFLOW_BUCKET}/mlflow-artifacts + - --host + - 0.0.0.0 + - --port + - "5000" + environment: + AWS_ACCESS_KEY_ID: ${RUSTFS_ACCESS_KEY} + AWS_SECRET_ACCESS_KEY: ${RUSTFS_SECRET_KEY} + MLFLOW_S3_ENDPOINT_URL: http://rustfs:9000 + AWS_DEFAULT_REGION: us-east-1 + depends_on: + create-bucket: + condition: service_completed_successfully + ports: + - "5000:5000" + volumes: + - mlflow-db:/mlflow + networks: + - mlflow + +networks: + mlflow: + +volumes: + rustfs-data: + mlflow-db: +``` + +The `create-bucket` service must run before the server starts because MLflow does not create the bucket. `MLFLOW_S3_ENDPOINT_URL` points the server's boto3 client at RustFS with path-style addressing. The SQLite database lives on a volume so experiment metadata survives restarts; production deployments should use a managed database backend instead. + +## 2. Validate and start the deployment + +Resolve the Compose file before starting containers: + +```bash +docker compose config +``` + +Start the services and wait for the tracking server: + +```bash +docker compose up -d +docker compose ps +``` + +The `create-bucket` service should exit with code `0`, and the MLflow UI should answer on `http://localhost:5000`: + +```bash +curl -sf http://localhost:5000/ >/dev/null && echo ready +``` + +Open the RustFS Console at `http://localhost:9001` to watch artifacts land in `my-bucket` during the next step. + +## 3. Log a training run + +Create the client script — it runs inside the MLflow image, which ships every dependency: + +```bash title="train_demo.py" {12} +import mlflow + +mlflow.set_tracking_uri("http://localhost:5000") +mlflow.set_experiment("rustfs-demo") + +with mlflow.start_run(run_name="rustfs-verify") as run: + mlflow.log_params({"model": "demo-regressor", "alpha": 0.5}) + for step in range(3): + mlflow.log_metric("rmse", 0.9 - step * 0.2, step=step) + with open("model-summary.txt", "w") as f: + f.write("demo model trained against RustFS artifact store\n") + mlflow.log_artifact("model-summary.txt", artifact_path="reports") + print("run_id:", run.info.run_id) + print("artifact_uri:", run.info.artifact_uri) +``` + +Copy the script into the running container and execute it with the service environment: + +```bash +docker compose cp train_demo.py mlflow:/tmp/train_demo.py +docker compose exec -w /tmp mlflow python train_demo.py +``` + +The `artifact_uri` prints as `s3://my-bucket/mlflow-artifacts///artifacts` — the artifact is uploaded straight to RustFS by the client. + +## 4. Verify artifacts in RustFS + +Read the artifact back through the tracking server, then list the same object in RustFS. Save the following as `verify.py`, replace `` with the identifier printed in step 3, and run it the same way: + +```bash title="verify.py" {5} +import mlflow + +mlflow.set_tracking_uri("http://localhost:5000") +client = mlflow.MlflowClient() +print([a.path for a in client.list_artifacts("", "reports")]) +path = client.download_artifacts("", "reports/model-summary.txt") +print(open(path).read()) +``` + +The download reads the object from RustFS through the tracking server. Then confirm the objects in the bucket: + +```bash +docker compose exec rustfs /usr/bin/rc ls local/my-bucket/mlflow-artifacts/ -r +``` + +The output should include the artifact object: + +```text +mlflow-artifacts/1/a1aece9243504f2680a52fba0c32765f/artifacts/reports/model-summary.txt +``` + +![MLflow artifacts stored in the RustFS Console](./images/rustfs-mlflow-artifacts.png) + +Runs, parameters, and metrics survive a server restart because they are stored in SQLite, while the artifacts stay in RustFS: + +```bash +docker compose restart mlflow +docker compose exec -w /tmp mlflow python verify.py +``` + +The `list_artifacts` call succeeds again against the restarted server, reading from the same objects in RustFS. + +## 5. Stop or reset the deployment + +Stop the containers while keeping all data: + +```bash +docker compose down +``` + +The RustFS volume keeps the artifact objects and the MLflow volume keeps the metadata database. To delete everything, including the artifacts in RustFS, add `--volumes`. + +## Troubleshooting + +### `ModuleNotFoundError: No module named 'boto3'` when logging artifacts + +The client performing `log_artifact` needs boto3 because it uploads directly to S3. Install it in the environment that runs the training script, or run the script inside the MLflow image as shown in step 3. + +### `AccessDenied` or connection errors during artifact upload + +Confirm that `MLFLOW_S3_ENDPOINT_URL` is set in the client environment — without it, boto3 sends requests to real AWS S3. The endpoint must be reachable from the machine running the training script; use `http://localhost:9000` outside the Compose network and `http://rustfs:9000` inside it. + +### The server fails to start with a bucket error + +The artifact bucket must exist before the server starts. Check the `create-bucket` service logs: + +```bash +docker compose logs create-bucket +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional MLflow operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [MLflow documentation](https://mlflow.org/docs/latest/) to add a model registry or move the metadata store to a managed database. diff --git a/content/fr/developer/integration/devops/gitea.md b/content/fr/developer/integration/devops/gitea.md new file mode 100644 index 00000000..239b6ad2 --- /dev/null +++ b/content/fr/developer/integration/devops/gitea.md @@ -0,0 +1,236 @@ +--- +title: "Gitea" +description: "Run Gitea with RustFS as the S3 storage backend for LFS objects and attachments, deployed with Docker Compose." +--- + +This guide connects [Gitea](https://github.com/go-gitea/gitea) — the self-hosted Git service — to **RustFS** through Gitea's `minio` storage type. You will start Gitea with Docker Compose, create a repository, push a Git LFS object, attach a file to an issue, and verify that both the LFS object and the attachment are stored in RustFS. The workflow was verified with `gitea/gitea:1.24.4` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with the Compose plugin and the `git` and `git-lfs` clients on your workstation. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Dev["Git + LFS client"] -->|"git push / git-lfs"| Gitea["Gitea :3000"] + Gitea -->|"LFS + attachments"| RustFS["RustFS :9000"] +``` + +Gitea stores the Git repository itself on its local disk, while the `minio` storage type routes large files — LFS objects, issue attachments, avatars, repository archives, packages, and Actions artifacts — to the `gitea-data` bucket in RustFS. Gitea creates the bucket on startup if it does not exist. + +## 1. Create the project files + +Create a working directory: + +```bash +mkdir rustfs-gitea +cd rustfs-gitea +``` + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +RUSTFS_ACCESS_KEY= +RUSTFS_SECRET_KEY= +``` + +Use dedicated credentials for the `gitea-data` bucket. Do not commit `.env` to source control. + +Create the Gitea configuration and replace both credential placeholders with the same values you set in `.env` — the global `minio` storage type applies to LFS, attachments, avatars, repository archives, packages, and Actions artifacts: + +```ini title="app.ini" +APP_NAME = RustFS Gitea +RUN_MODE = prod +WORK_PATH = /data/gitea + +[server] +DOMAIN = localhost +ROOT_URL = http://localhost:3000/ +HTTP_PORT = 3000 +LFS_START_SERVER = true + +[database] +DB_TYPE = sqlite3 +PATH = /data/gitea/gitea.db + +[storage] +STORAGE_TYPE = minio +MINIO_ENDPOINT = rustfs:9000 +MINIO_ACCESS_KEY_ID = +MINIO_SECRET_ACCESS_KEY = +MINIO_BUCKET = gitea-data +MINIO_LOCATION = us-east-1 +MINIO_USE_SSL = false + +[log] +MODE = console +LEVEL = info + +[security] +INSTALL_LOCK = true +SECRET_KEY = change-me-to-a-random-string +``` + +`LFS_START_SERVER` enables the Git LFS HTTP API. `MINIO_ENDPOINT` uses the Compose-internal hostname `rustfs`; the endpoint is reached with path-style requests by default. + +Create the Compose file: + +```yaml title="compose.yaml" +services: + rustfs: + image: rustfs/rustfs-x86-musl:v2.3.1 + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + RUSTFS_VOLUMES: /data + RUSTFS_ADDRESS: ":9000" + RUSTFS_CONSOLE_ADDRESS: ":9001" + RUSTFS_CONSOLE_ENABLE: "true" + volumes: + - rustfs-data:/data + ports: + - "9000:9000" + - "9001:9001" + healthcheck: + test: ["CMD", "curl", "-sf", "http://127.0.0.1:9000/health"] + interval: 10s + timeout: 5s + retries: 6 + start_period: 10s + networks: + - gitea + + gitea: + image: gitea/gitea:1.24.4 + depends_on: + rustfs: + condition: service_healthy + environment: + USER_UID: "1000" + USER_GID: "1000" + volumes: + - gitea-data:/data + - ./app.ini:/data/gitea/conf/app.ini:ro + ports: + - "3000:3000" + networks: + - gitea + +networks: + gitea: + +volumes: + rustfs-data: + gitea-data: +``` + +The volume-backed `/data` directory keeps the SQLite database and the Git repositories across container restarts, while LFS objects and attachments live in RustFS. + +## 2. Validate and start the deployment + +Resolve the Compose file before starting containers: + +```bash +docker compose config +``` + +Start the services: + +```bash +docker compose up -d +docker compose ps +``` + +Watch the Gitea log until every storage backend reports the Minio type: + +```bash +docker compose logs gitea | grep "Initialising" +``` + +The output should list `Attachment`, `Avatar`, `LFS`, and the remaining storage sections, each followed by a `Creating Minio storage at rustfs:9000:gitea-data` line. + +Open `http://localhost:3000` and create the administrator account, then open the RustFS Console at `http://localhost:9001` — the `gitea-data` bucket appears after the first storage operation. + +## 3. Push a Git LFS object + +Create a repository named `rustfs-demo` in the Gitea web UI, then push an LFS-tracked file from your workstation: + +```bash +mkdir lfs-demo && cd lfs-demo +git init +git config user.email you@example.com +git config user.name you +git lfs install +git lfs track "*.bin" +git add .gitattributes +dd if=/dev/urandom of=dataset.bin bs=1M count=8 +git add dataset.bin +git commit -m "add LFS dataset" +git remote add origin http://localhost:3000//rustfs-demo.git +git push origin main +``` + +`git push` uploads the LFS object through the Gitea LFS API, which writes it to RustFS. Clone the repository into a second directory and run `git lfs pull` — the downloaded `dataset.bin` must be byte-identical to the original: + +```bash +sha256sum dataset.bin +cd ../lfs-demo-clone && git lfs pull && sha256sum dataset.bin +``` + +Both checksums match because both clients read the object from RustFS. + +## 4. Attach a file to an issue + +Open the `rustfs-demo` repository, create an issue, and attach a small text file through the issue form. Gitea stores the upload as `attachments//` in the `gitea-data` bucket and serves downloads through `/attachments/`. + +## 5. Verify objects in RustFS + +List the bucket with the [`rc` client](https://github.com/rustfs/cli): + +```bash +docker compose exec rustfs /usr/bin/rc ls local/gitea-data/ -r +``` + +The output should include the LFS object under `lfs/` and the attachment under `attachments/`: + +```text +attachments/9/2/92fdd48d-531c-4cba-8b3f-4e2004a10fc7 +lfs/37/76/6ddfc07e803de58a69328db9a58a07cf7080ddde55c155a7531bc650a000 +``` + +The LFS object key is the SHA-256 content hash used by the Git LFS protocol. + +![Gitea LFS and attachment objects in the RustFS Console](./images/rustfs-gitea-objects.png) + +## 6. Stop or reset the deployment + +Stop the containers while keeping all data: + +```bash +docker compose down +``` + +The RustFS volume keeps the `gitea-data` bucket, so LFS objects and attachments survive a restart. To delete everything, including the objects in RustFS, add `--volumes`. + +## Troubleshooting + +### The Gitea install page appears instead of the login page + +The configuration file must exist at `/data/gitea/conf/app.ini` inside the container. If the mount path is wrong, Gitea starts with defaults and shows the installation wizard. Mount the file as shown in the Compose example and restart. + +### Push fails with an LFS or 403 error + +Confirm the credentials in `app.ini` match the RustFS credentials and that the `rustfs` hostname resolves inside the Compose network: + +```bash +docker compose logs gitea | grep -i minio +``` + +### Objects land in local storage instead of RustFS + +The `GITEA__storage__STORAGE_TYPE: minio` environment variable and the `[storage]` section of `app.ini` must agree. After changing either, restart Gitea and check the `Initialising` log lines again. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Gitea storage targets. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Gitea storage documentation](https://docs.gitea.com/administration/storage-configurations) to move packages, Actions artifacts, or individual storage sections to separate buckets. diff --git a/content/fr/developer/integration/devops/images/rustfs-gitea-objects.png b/content/fr/developer/integration/devops/images/rustfs-gitea-objects.png new file mode 100644 index 00000000..682f280c Binary files /dev/null and b/content/fr/developer/integration/devops/images/rustfs-gitea-objects.png differ diff --git a/content/fr/developer/integration/devops/index.md b/content/fr/developer/integration/devops/index.md index 1b4d0f17..0f991539 100644 --- a/content/fr/developer/integration/devops/index.md +++ b/content/fr/developer/integration/devops/index.md @@ -8,6 +8,7 @@ Utilisez **RustFS** comme couche de stockage objet pour les plateformes DevOps e ## Plateformes et outils - [Elasticsearch](./elasticsearch.md) +- [Gitea](./gitea.md) - [Terraform](./terraform.md) Conservez les artefacts, l'état et les données de télémétrie dans des buckets dédiés et limitez les identifiants aux opérations de bucket requises. diff --git a/content/fr/developer/integration/devops/meta.json b/content/fr/developer/integration/devops/meta.json index a032acbd..61ad9fb8 100644 --- a/content/fr/developer/integration/devops/meta.json +++ b/content/fr/developer/integration/devops/meta.json @@ -2,6 +2,7 @@ "title": "DevOps", "pages": [ "elasticsearch", + "gitea", "terraform" ] } diff --git a/content/fr/developer/integration/index.md b/content/fr/developer/integration/index.md index 986f76de..0af98703 100644 --- a/content/fr/developer/integration/index.md +++ b/content/fr/developer/integration/index.md @@ -8,11 +8,11 @@ Utilisez cette section pour connecter **RustFS** à des plateformes d'infrastruc ## Integration categories - [Reverse Proxy](./reverse-proxy/index.md) couvre Nginx, Traefik, Caddy et HAProxy. -- [Backup](./backup/index.md) couvre Restic. +- [Backup](./backup/index.md) couvre Restic et Longhorn. - [Analyse de données](./big-data/index.md) couvre Iceberg. - [Observabilité](./observability/index.md) couvre OpenObserve. - [Autres](./others/index.md) couvre le SDK communautaire capo pour Python. - [Registre](./registry/index.md) couvre Harbor. -- [DevOps](./devops/index.md) couvre Elasticsearch et Terraform. +- [DevOps](./devops/index.md) couvre Elasticsearch, Gitea et Terraform. Chaque guide indique le point de terminaison RustFS et les exigences d'adressage à utiliser lors de la configuration du système intégré. \ No newline at end of file diff --git a/content/fr/developer/integration/observability/images/rustfs-thanos-blocks.png b/content/fr/developer/integration/observability/images/rustfs-thanos-blocks.png new file mode 100644 index 00000000..188cb9cf Binary files /dev/null and b/content/fr/developer/integration/observability/images/rustfs-thanos-blocks.png differ diff --git a/content/fr/developer/integration/observability/index.md b/content/fr/developer/integration/observability/index.md index cd3d854e..2094641a 100644 --- a/content/fr/developer/integration/observability/index.md +++ b/content/fr/developer/integration/observability/index.md @@ -10,5 +10,6 @@ Utilisez **RustFS** comme couche de stockage objet pour les plateformes d'observ - [OpenObserve](./openobserve.md) - [Loki](./loki.md) - [Tempo](./tempo.md) +- [Thanos](./thanos.md) Conservez les données de télémétrie dans un bucket dédié et limitez les identifiants aux opérations de bucket requises. diff --git a/content/fr/developer/integration/observability/meta.json b/content/fr/developer/integration/observability/meta.json index c9fe1865..1709c8c7 100644 --- a/content/fr/developer/integration/observability/meta.json +++ b/content/fr/developer/integration/observability/meta.json @@ -3,6 +3,7 @@ "pages": [ "openobserve", "loki", - "tempo" + "tempo", + "thanos" ] } diff --git a/content/fr/developer/integration/observability/thanos.md b/content/fr/developer/integration/observability/thanos.md new file mode 100644 index 00000000..856c31db --- /dev/null +++ b/content/fr/developer/integration/observability/thanos.md @@ -0,0 +1,311 @@ +--- +title: "Thanos" +description: "Run Thanos with RustFS as the S3 object storage backend for Prometheus blocks, deployed with Docker Compose." +--- + +This guide connects [Thanos](https://github.com/thanos-io/thanos) — the highly available Prometheus setup with long-term storage — to **RustFS** as its object store. You will run Prometheus with a Thanos sidecar that uploads TSDB blocks to RustFS, then query the historical data back through a Store Gateway and a Query frontend. The workflow was verified with `thanosio/thanos:v0.37.2`, `prom/prometheus:v2.53.1`, and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with the Compose plugin. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Prom["Prometheus :9090"] -->|"blocks"| Sidecar["Thanos sidecar"] + Sidecar -->|"upload"| RustFS["RustFS :9000"] + Store["Store Gateway"] -->|"download"| RustFS + Query["Thanos Query"] -->|gRPC| Sidecar + Query -->|gRPC| Store +``` + +The sidecar watches the Prometheus TSDB directory and uploads every two-hour block to the `thanos-data` bucket in RustFS. The Store Gateway reads the same bucket and answers queries about historical blocks, so Query resolves both live data through the sidecar and old data through the Store Gateway. + +## 1. Create the project files + +Create a working directory: + +```bash +mkdir rustfs-thanos +cd rustfs-thanos +``` + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +RUSTFS_ACCESS_KEY= +RUSTFS_SECRET_KEY= +``` + +Use dedicated credentials for the `thanos-data` bucket. Do not commit `.env` to source control. + +Create the Prometheus configuration with an external label — Thanos requires it to deduplicate blocks: + +```yaml title="prometheus.yml" +global: + scrape_interval: 5s + external_labels: + monitor: rustfs-demo + +scrape_configs: + - job_name: prometheus + static_configs: + - targets: ["localhost:9090"] + - job_name: rustfs + metrics_path: /metrics + static_configs: + - targets: ["rustfs:9000"] +``` + +Create the Thanos object store configuration: + +```yaml title="bucket.yml" +type: S3 +config: + bucket: thanos-data + endpoint: rustfs:9000 + access_key: ${RUSTFS_ACCESS_KEY} + secret_key: ${RUSTFS_SECRET_KEY} + insecure: true +``` + +Thanos does not interpolate `.env` files itself. Before starting the stack, replace the placeholders with the same values you set in `.env`: + +```bash +sed -i.bak "s|\${RUSTFS_ACCESS_KEY}|$(grep RUSTFS_ACCESS_KEY .env | cut -d= -f2)|;s|\${RUSTFS_SECRET_KEY}|$(grep RUSTFS_SECRET_KEY .env | cut -d= -f2)|" bucket.yml +``` + +Create the Compose file: + +```yaml title="compose.yaml" +services: + rustfs: + image: rustfs/rustfs-x86-musl:v2.3.1 + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + RUSTFS_VOLUMES: /data + RUSTFS_ADDRESS: ":9000" + RUSTFS_CONSOLE_ADDRESS: ":9001" + RUSTFS_CONSOLE_ENABLE: "true" + volumes: + - rustfs-data:/data + ports: + - "9000:9000" + - "9001:9001" + healthcheck: + test: ["CMD", "curl", "-sf", "http://127.0.0.1:9000/health"] + interval: 10s + timeout: 5s + retries: 6 + start_period: 10s + networks: + - thanos + + create-bucket: + image: rustfs/rc:latest + depends_on: + rustfs: + condition: service_healthy + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + entrypoint: + - /bin/sh + - -c + - | + /usr/bin/rc alias set rustfs http://rustfs:9000 "$${RUSTFS_ACCESS_KEY}" "$${RUSTFS_SECRET_KEY}" + /usr/bin/rc mb --ignore-existing rustfs/thanos-data + networks: + - thanos + + prometheus: + image: prom/prometheus:v2.53.1 + command: + - --config.file=/etc/prometheus/prometheus.yml + - --storage.tsdb.path=/prometheus + - --storage.tsdb.min-block-duration=2h + - --storage.tsdb.max-block-duration=2h + - --web.enable-lifecycle + volumes: + - ./prometheus.yml:/etc/prometheus/prometheus.yml:ro + - prom-data:/prometheus + ports: + - "9090:9090" + networks: + - thanos + + sidecar: + image: thanosio/thanos:v0.37.2 + command: + - sidecar + - --tsdb.path=/prometheus + - --prometheus.url=http://prometheus:9090 + - --objstore.config-file=/etc/thanos/bucket.yml + volumes: + - ./bucket.yml:/etc/thanos/bucket.yml:ro + - prom-data:/prometheus + depends_on: + create-bucket: + condition: service_completed_successfully + networks: + - thanos + + store: + image: thanosio/thanos:v0.37.2 + command: + - store + - --objstore.config-file=/etc/thanos/bucket.yml + - --data-dir=/data + volumes: + - ./bucket.yml:/etc/thanos/bucket.yml:ro + - store-data:/data + depends_on: + create-bucket: + condition: service_completed_successfully + networks: + - thanos + + query: + image: thanosio/thanos:v0.37.2 + command: + - query + - --http-address=0.0.0.0:9090 + - --store=sidecar:10901 + - --store=store:10901 + ports: + - "9091:9090" + depends_on: + - sidecar + - store + networks: + - thanos + +networks: + thanos: + +volumes: + rustfs-data: + prom-data: + store-data: +``` + +The `--storage.tsdb.min-block-duration` and `--storage.tsdb.max-block-duration` flags disable Prometheus compaction. The sidecar refuses to ship blocks from a compacting TSDB because the local blocks would no longer match the uploaded ones. + +## 2. Validate and start the deployment + +Resolve the Compose file before starting containers: + +```bash +docker compose config +``` + +Start the stack and wait until the sidecar reports itself ready: + +```bash +docker compose up -d +docker compose logs sidecar | grep -m1 "status=ready" +``` + +Check that the sidecar picked up the Prometheus external labels: + +```bash +docker compose logs sidecar | grep "external labels" +``` + +The Thanos Query UI answers on `http://localhost:9091`, and the RustFS Console runs at `http://localhost:9001`. + +## 3. Upload a block to RustFS + +The sidecar uploads a block when Prometheus compacts one, which happens at a two-hour block boundary. To produce a block immediately, snapshot the TSDB through the admin API — with compaction disabled, the sidecar ships the head-block snapshot directly: + +```bash +curl -s -XPOST http://localhost:9090/api/v1/admin/tsdb/snapshot | head -c 200 +``` + +Wait for the upload, then check the shipper state inside Prometheus: + +```bash +sleep 60 +docker compose exec prometheus cat /prometheus/thanos.shipper.json +``` + +The `uploaded` list should contain a block ID: + +```json +{ + "version": 1, + "uploaded": [ + "01M31EPTZC5E0SETZTP0SPFY79" + ] +} +``` + +## 4. Query historical data from RustFS + +The Store Gateway periodically syncs the bucket. Confirm it downloaded the uploaded block: + +```bash +docker compose logs store | grep "loaded new block" +``` + +Query a series through the Query frontend over the block's time range: + +```bash +START=$(date -u -d '2 hours ago' +%s) +END=$(date -u +%s) +curl -s "http://localhost:9091/api/v1/query_range?query=up%7Bjob%3D%22prometheus%22%7D&start=$START&end=$END&step=30" | head -c 300 +``` + +The Store Gateway serves the response from the blocks it downloaded from RustFS, while the sidecar answers for the live head — both paths resolve through the same Query endpoint. + +## 5. Verify objects in RustFS + +List the bucket: + +```bash +docker compose exec rustfs /usr/bin/rc ls local/thanos-data/ -r +``` + +Each block is stored as three objects — the chunk files, the index, and `meta.json`: + +```text +01M31EPTZC5E0SETZTP0SPFY79/chunks/000001 +01M31EPTZC5E0SETZTP0SPFY79/index +01M31EPTZC5E0SETZTP0SPFY79/meta.json +``` + +![Thanos blocks stored in the RustFS Console](./images/rustfs-thanos-blocks.png) + +## 6. Stop or reset the deployment + +Stop the containers while keeping all data: + +```bash +docker compose down +``` + +The RustFS volume keeps the uploaded blocks, so the Store Gateway serves historical queries again after a restart. To delete everything, including the blocks in RustFS, add `--volumes`. + +## Troubleshooting + +### The sidecar logs `Compaction needs to be disabled` + +Prometheus must run with `--storage.tsdb.min-block-duration` equal to `--storage.tsdb.max-block-duration` — set both to `2h` as shown in the Compose file. Otherwise the sidecar cannot guarantee that local blocks stay unchanged and refuses to upload. + +### `The specified bucket does not exist` + +Thanos does not create buckets. Check that the `create-bucket` service completed successfully: + +```bash +docker compose logs create-bucket +``` + +### Queries return no historical data + +Confirm that the Store Gateway has loaded at least one block (`docker compose logs store | grep "loaded new block"`) and that your query time range falls inside the uploaded block's window — check the block's `meta.json` in the RustFS Console for `minTime` and `maxTime`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Thanos components. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Thanos documentation](https://thanos.io/tip/thanos/getting-started.md) to add Compactor, Ruler, or Receive for a production topology. diff --git a/content/ja/developer/integration/backup/images/rustfs-longhorn-backups.png b/content/ja/developer/integration/backup/images/rustfs-longhorn-backups.png new file mode 100644 index 00000000..899b1c8c Binary files /dev/null and b/content/ja/developer/integration/backup/images/rustfs-longhorn-backups.png differ diff --git a/content/ja/developer/integration/backup/index.md b/content/ja/developer/integration/backup/index.md index 37cd3e8f..63f7958a 100644 --- a/content/ja/developer/integration/backup/index.md +++ b/content/ja/developer/integration/backup/index.md @@ -8,5 +8,6 @@ description: "S3 互換のオブジェクトストレージ経由でバックア ## システム - [Restic](./restic.md) +- [Longhorn](./longhorn.md) バックアップジョブは専用のバケットとプレフィックスにまとめ、必要なバケット操作だけに絞った認証情報を使用します。 \ No newline at end of file diff --git a/content/ja/developer/integration/backup/longhorn.md b/content/ja/developer/integration/backup/longhorn.md new file mode 100644 index 00000000..497f43cd --- /dev/null +++ b/content/ja/developer/integration/backup/longhorn.md @@ -0,0 +1,293 @@ +--- +title: "Longhorn" +description: "Configure Longhorn to store Kubernetes volume backups in RustFS through its S3 backup target." +--- + +This guide connects [Longhorn](https://github.com/longhorn/longhorn) — the distributed block storage system for Kubernetes — to **RustFS** as its S3 backup target. You will configure the backup target, back up a volume that contains data, delete the volume, restore it from RustFS, and verify the data. The workflow was verified with Longhorn 1.9.0 on k3s (Kubernetes 1.30) and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need a Kubernetes cluster with Longhorn installed and `kubectl` access to it. This guide is intended for integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + App["Workload pod"] -->|"writes"| Vol["Longhorn volume"] + Vol -->|"snapshot"| Backup["Backup engine"] + Backup -->|"blocks + config"| RustFS["RustFS :9000"] +``` + +Longhorn stores backups in the `backupstore/volumes/` prefix of the target bucket as content-addressed blocks plus a small volume configuration file. Restores read those objects back into a new volume on any node that can reach RustFS. + +## 1. Create the backup bucket + +Create a dedicated bucket with the [`rc` client](https://github.com/rustfs/cli), using your own endpoint and credentials: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/longhorn-backups +``` + +Longhorn does not create the bucket, so this step must complete before the first backup. + +## 2. Configure the backup target + +Store the RustFS credentials in a secret in the `longhorn-system` namespace. `AWS_ENDPOINTS` must be an endpoint every node can reach — use the node IP or an internal load balancer address, not a port forward from your workstation: + +```bash +kubectl -n longhorn-system create secret generic rustfs-s3-secret \ + --from-literal=AWS_ACCESS_KEY_ID= \ + --from-literal=AWS_SECRET_ACCESS_KEY= \ + --from-literal=AWS_ENDPOINTS=http://:9000 +``` + +Longhorn 1.9 manages backup targets through the `BackupTarget` resource. Patch the `default` target with the RustFS bucket: + +```bash +kubectl -n longhorn-system patch backupTarget default --type merge -p ' +spec: + backupTargetURL: s3://longhorn-backups@us-east-1/ + credentialSecret: rustfs-s3-secret + pollInterval: 5m' +``` + +The `@us-east-1` segment is the region annotation in the S3 URL format; it does not need to match a real deployment region. + +Wait for the target to become available — it confirms that Longhorn reached the bucket through the secret: + +```bash +kubectl -n longhorn-system get backupTarget default +``` + +```text +NAME URL CREDENTIAL AVAILABLE LASTSYNCEDAT +default s3://longhorn-backups@us-east-1/ rustfs-s3-secret true 2026-09-21T08:36:55Z +``` + +## 3. Write data and create a backup + +Create a test volume with data in it. The following manifest creates a 1 GiB PVC and a pod that writes a marker file: + +```yaml title="demo.yaml" +apiVersion: v1 +kind: PersistentVolumeClaim +metadata: + name: demo-vol +spec: + accessModes: [ReadWriteOnce] + storageClassName: longhorn + resources: + requests: + storage: 1Gi +--- +apiVersion: v1 +kind: Pod +metadata: + name: demo-app +spec: + volumes: + - name: data + persistentVolumeClaim: + claimName: demo-vol + containers: + - name: app + image: busybox:1.36 + command: ["sh", "-c", "echo 'longhorn rustfs demo' > /data/hello.txt && sleep 3600"] + volumeMounts: + - name: data + mountPath: /data +``` + +Apply it and wait for the pod to run: + +```bash +kubectl apply -f demo.yaml +kubectl get pod demo-app +``` + +Create a snapshot and back it up. You can do this in the Longhorn UI (Volume → Snapshot → Backup) or declaratively: + +```bash +VOLUME=$(kubectl get pvc demo-vol -o jsonpath='{.spec.volumeName}') +kubectl -n longhorn-system apply -f - <" +EOF +``` + +The restore runs when the volume is first attached. Create a PVC bound to the restored volume through a static PV, then mount it: + +```bash +kubectl apply -f - < --type merge -p ' +spec: + disks: + default: + path: /var/lib/longhorn + allowScheduling: true' +``` + +### `failed to create backup ... missing input parameter` + +The backup ran before the default backup target was configured. Complete step 2, confirm `AVAILABLE` is `true`, and create the backup again. + +### The backup target never becomes available + +The secret must exist before the target syncs, and `AWS_ENDPOINTS` must be reachable from the nodes themselves. Check the `longhorn-manager` logs for S3 errors: + +```bash +kubectl -n longhorn-system logs -l app=longhorn-manager | grep -i s3 | tail +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional backup targets. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Longhorn backup documentation](https://longhorn.io/docs/1.9.0/backups-and-restore/) to configure recurring backup jobs and scheduled snapshots. diff --git a/content/ja/developer/integration/backup/meta.json b/content/ja/developer/integration/backup/meta.json index 123f4744..6cd71cbf 100644 --- a/content/ja/developer/integration/backup/meta.json +++ b/content/ja/developer/integration/backup/meta.json @@ -1,6 +1,7 @@ { "title": "バックアップ", "pages": [ - "restic" + "restic", + "longhorn" ] -} \ No newline at end of file +} diff --git a/content/ja/developer/integration/big-data/images/rustfs-mlflow-artifacts.png b/content/ja/developer/integration/big-data/images/rustfs-mlflow-artifacts.png new file mode 100644 index 00000000..545f5bf2 Binary files /dev/null and b/content/ja/developer/integration/big-data/images/rustfs-mlflow-artifacts.png differ diff --git a/content/ja/developer/integration/big-data/index.md b/content/ja/developer/integration/big-data/index.md index 22510c10..a64033c9 100644 --- a/content/ja/developer/integration/big-data/index.md +++ b/content/ja/developer/integration/big-data/index.md @@ -10,6 +10,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) - [Milvus](./milvus.md) +- [MLflow](./mlflow.md) - [DuckDB](./duckdb.md) - [InfluxDB](./influxdb.md) - [Spark](./spark.md) diff --git a/content/ja/developer/integration/big-data/meta.json b/content/ja/developer/integration/big-data/meta.json index b2fbd258..ebbb0f87 100644 --- a/content/ja/developer/integration/big-data/meta.json +++ b/content/ja/developer/integration/big-data/meta.json @@ -4,6 +4,7 @@ "iceberg", "pyiceberg", "milvus", + "mlflow", "duckdb", "influxdb", "spark", diff --git a/content/ja/developer/integration/big-data/mlflow.md b/content/ja/developer/integration/big-data/mlflow.md new file mode 100644 index 00000000..46d3205e --- /dev/null +++ b/content/ja/developer/integration/big-data/mlflow.md @@ -0,0 +1,244 @@ +--- +title: "MLflow" +description: "Run MLflow with RustFS as the S3 artifact store for experiment tracking, deployed with Docker Compose." +--- + +This guide connects [MLflow](https://github.com/mlflow/mlflow) — the experiment tracking and model registry platform — to **RustFS** as its S3 artifact store. You will start the MLflow tracking server with Docker Compose, log parameters, metrics, and artifacts from a training run, and verify that the artifacts are stored in RustFS. The workflow was verified with `ghcr.io/mlflow/mlflow:v2.22.1` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with the Compose plugin. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["Training script"] -->|"runs + metrics"| Server["MLflow server :5000"] + Client -->|"artifacts"| RustFS["RustFS :9000"] + Server -->|"metadata"| DB["SQLite"] +``` + +The tracking server keeps experiment and run metadata in SQLite and stores artifacts — model files, plots, reports — directly in RustFS through the `s3://` artifact root. The client needs the same RustFS credentials because it uploads artifacts itself, using boto3 and the `MLFLOW_S3_ENDPOINT_URL` setting. + +## 1. Create the project files + +Create a working directory: + +```bash +mkdir rustfs-mlflow +cd rustfs-mlflow +``` + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +RUSTFS_ACCESS_KEY= +RUSTFS_SECRET_KEY= +MLFLOW_BUCKET=my-bucket +``` + +Use dedicated credentials for the artifact bucket. Do not commit `.env` to source control. + +Create the Compose file: + +```yaml title="compose.yaml" +services: + rustfs: + image: rustfs/rustfs-x86-musl:v2.3.1 + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + RUSTFS_VOLUMES: /data + RUSTFS_ADDRESS: ":9000" + RUSTFS_CONSOLE_ADDRESS: ":9001" + RUSTFS_CONSOLE_ENABLE: "true" + volumes: + - rustfs-data:/data + ports: + - "9000:9000" + - "9001:9001" + healthcheck: + test: ["CMD", "curl", "-sf", "http://127.0.0.1:9000/health"] + interval: 10s + timeout: 5s + retries: 6 + start_period: 10s + networks: + - mlflow + + create-bucket: + image: rustfs/rc:latest + depends_on: + rustfs: + condition: service_healthy + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + MLFLOW_BUCKET: ${MLFLOW_BUCKET} + entrypoint: + - /bin/sh + - -c + - | + /usr/bin/rc alias set rustfs http://rustfs:9000 "$${RUSTFS_ACCESS_KEY}" "$${RUSTFS_SECRET_KEY}" + /usr/bin/rc mb --ignore-existing rustfs/$${MLFLOW_BUCKET} + networks: + - mlflow + + mlflow: + image: ghcr.io/mlflow/mlflow:v2.22.1 + command: + - server + - --backend-store-uri + - sqlite:////mlflow/mlflow.db + - --default-artifact-root + - s3://${MLFLOW_BUCKET}/mlflow-artifacts + - --host + - 0.0.0.0 + - --port + - "5000" + environment: + AWS_ACCESS_KEY_ID: ${RUSTFS_ACCESS_KEY} + AWS_SECRET_ACCESS_KEY: ${RUSTFS_SECRET_KEY} + MLFLOW_S3_ENDPOINT_URL: http://rustfs:9000 + AWS_DEFAULT_REGION: us-east-1 + depends_on: + create-bucket: + condition: service_completed_successfully + ports: + - "5000:5000" + volumes: + - mlflow-db:/mlflow + networks: + - mlflow + +networks: + mlflow: + +volumes: + rustfs-data: + mlflow-db: +``` + +The `create-bucket` service must run before the server starts because MLflow does not create the bucket. `MLFLOW_S3_ENDPOINT_URL` points the server's boto3 client at RustFS with path-style addressing. The SQLite database lives on a volume so experiment metadata survives restarts; production deployments should use a managed database backend instead. + +## 2. Validate and start the deployment + +Resolve the Compose file before starting containers: + +```bash +docker compose config +``` + +Start the services and wait for the tracking server: + +```bash +docker compose up -d +docker compose ps +``` + +The `create-bucket` service should exit with code `0`, and the MLflow UI should answer on `http://localhost:5000`: + +```bash +curl -sf http://localhost:5000/ >/dev/null && echo ready +``` + +Open the RustFS Console at `http://localhost:9001` to watch artifacts land in `my-bucket` during the next step. + +## 3. Log a training run + +Create the client script — it runs inside the MLflow image, which ships every dependency: + +```bash title="train_demo.py" {12} +import mlflow + +mlflow.set_tracking_uri("http://localhost:5000") +mlflow.set_experiment("rustfs-demo") + +with mlflow.start_run(run_name="rustfs-verify") as run: + mlflow.log_params({"model": "demo-regressor", "alpha": 0.5}) + for step in range(3): + mlflow.log_metric("rmse", 0.9 - step * 0.2, step=step) + with open("model-summary.txt", "w") as f: + f.write("demo model trained against RustFS artifact store\n") + mlflow.log_artifact("model-summary.txt", artifact_path="reports") + print("run_id:", run.info.run_id) + print("artifact_uri:", run.info.artifact_uri) +``` + +Copy the script into the running container and execute it with the service environment: + +```bash +docker compose cp train_demo.py mlflow:/tmp/train_demo.py +docker compose exec -w /tmp mlflow python train_demo.py +``` + +The `artifact_uri` prints as `s3://my-bucket/mlflow-artifacts///artifacts` — the artifact is uploaded straight to RustFS by the client. + +## 4. Verify artifacts in RustFS + +Read the artifact back through the tracking server, then list the same object in RustFS. Save the following as `verify.py`, replace `` with the identifier printed in step 3, and run it the same way: + +```bash title="verify.py" {5} +import mlflow + +mlflow.set_tracking_uri("http://localhost:5000") +client = mlflow.MlflowClient() +print([a.path for a in client.list_artifacts("", "reports")]) +path = client.download_artifacts("", "reports/model-summary.txt") +print(open(path).read()) +``` + +The download reads the object from RustFS through the tracking server. Then confirm the objects in the bucket: + +```bash +docker compose exec rustfs /usr/bin/rc ls local/my-bucket/mlflow-artifacts/ -r +``` + +The output should include the artifact object: + +```text +mlflow-artifacts/1/a1aece9243504f2680a52fba0c32765f/artifacts/reports/model-summary.txt +``` + +![MLflow artifacts stored in the RustFS Console](./images/rustfs-mlflow-artifacts.png) + +Runs, parameters, and metrics survive a server restart because they are stored in SQLite, while the artifacts stay in RustFS: + +```bash +docker compose restart mlflow +docker compose exec -w /tmp mlflow python verify.py +``` + +The `list_artifacts` call succeeds again against the restarted server, reading from the same objects in RustFS. + +## 5. Stop or reset the deployment + +Stop the containers while keeping all data: + +```bash +docker compose down +``` + +The RustFS volume keeps the artifact objects and the MLflow volume keeps the metadata database. To delete everything, including the artifacts in RustFS, add `--volumes`. + +## Troubleshooting + +### `ModuleNotFoundError: No module named 'boto3'` when logging artifacts + +The client performing `log_artifact` needs boto3 because it uploads directly to S3. Install it in the environment that runs the training script, or run the script inside the MLflow image as shown in step 3. + +### `AccessDenied` or connection errors during artifact upload + +Confirm that `MLFLOW_S3_ENDPOINT_URL` is set in the client environment — without it, boto3 sends requests to real AWS S3. The endpoint must be reachable from the machine running the training script; use `http://localhost:9000` outside the Compose network and `http://rustfs:9000` inside it. + +### The server fails to start with a bucket error + +The artifact bucket must exist before the server starts. Check the `create-bucket` service logs: + +```bash +docker compose logs create-bucket +``` + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional MLflow operations. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [MLflow documentation](https://mlflow.org/docs/latest/) to add a model registry or move the metadata store to a managed database. diff --git a/content/ja/developer/integration/devops/gitea.md b/content/ja/developer/integration/devops/gitea.md new file mode 100644 index 00000000..239b6ad2 --- /dev/null +++ b/content/ja/developer/integration/devops/gitea.md @@ -0,0 +1,236 @@ +--- +title: "Gitea" +description: "Run Gitea with RustFS as the S3 storage backend for LFS objects and attachments, deployed with Docker Compose." +--- + +This guide connects [Gitea](https://github.com/go-gitea/gitea) — the self-hosted Git service — to **RustFS** through Gitea's `minio` storage type. You will start Gitea with Docker Compose, create a repository, push a Git LFS object, attach a file to an issue, and verify that both the LFS object and the attachment are stored in RustFS. The workflow was verified with `gitea/gitea:1.24.4` and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with the Compose plugin and the `git` and `git-lfs` clients on your workstation. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Dev["Git + LFS client"] -->|"git push / git-lfs"| Gitea["Gitea :3000"] + Gitea -->|"LFS + attachments"| RustFS["RustFS :9000"] +``` + +Gitea stores the Git repository itself on its local disk, while the `minio` storage type routes large files — LFS objects, issue attachments, avatars, repository archives, packages, and Actions artifacts — to the `gitea-data` bucket in RustFS. Gitea creates the bucket on startup if it does not exist. + +## 1. Create the project files + +Create a working directory: + +```bash +mkdir rustfs-gitea +cd rustfs-gitea +``` + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +RUSTFS_ACCESS_KEY= +RUSTFS_SECRET_KEY= +``` + +Use dedicated credentials for the `gitea-data` bucket. Do not commit `.env` to source control. + +Create the Gitea configuration and replace both credential placeholders with the same values you set in `.env` — the global `minio` storage type applies to LFS, attachments, avatars, repository archives, packages, and Actions artifacts: + +```ini title="app.ini" +APP_NAME = RustFS Gitea +RUN_MODE = prod +WORK_PATH = /data/gitea + +[server] +DOMAIN = localhost +ROOT_URL = http://localhost:3000/ +HTTP_PORT = 3000 +LFS_START_SERVER = true + +[database] +DB_TYPE = sqlite3 +PATH = /data/gitea/gitea.db + +[storage] +STORAGE_TYPE = minio +MINIO_ENDPOINT = rustfs:9000 +MINIO_ACCESS_KEY_ID = +MINIO_SECRET_ACCESS_KEY = +MINIO_BUCKET = gitea-data +MINIO_LOCATION = us-east-1 +MINIO_USE_SSL = false + +[log] +MODE = console +LEVEL = info + +[security] +INSTALL_LOCK = true +SECRET_KEY = change-me-to-a-random-string +``` + +`LFS_START_SERVER` enables the Git LFS HTTP API. `MINIO_ENDPOINT` uses the Compose-internal hostname `rustfs`; the endpoint is reached with path-style requests by default. + +Create the Compose file: + +```yaml title="compose.yaml" +services: + rustfs: + image: rustfs/rustfs-x86-musl:v2.3.1 + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + RUSTFS_VOLUMES: /data + RUSTFS_ADDRESS: ":9000" + RUSTFS_CONSOLE_ADDRESS: ":9001" + RUSTFS_CONSOLE_ENABLE: "true" + volumes: + - rustfs-data:/data + ports: + - "9000:9000" + - "9001:9001" + healthcheck: + test: ["CMD", "curl", "-sf", "http://127.0.0.1:9000/health"] + interval: 10s + timeout: 5s + retries: 6 + start_period: 10s + networks: + - gitea + + gitea: + image: gitea/gitea:1.24.4 + depends_on: + rustfs: + condition: service_healthy + environment: + USER_UID: "1000" + USER_GID: "1000" + volumes: + - gitea-data:/data + - ./app.ini:/data/gitea/conf/app.ini:ro + ports: + - "3000:3000" + networks: + - gitea + +networks: + gitea: + +volumes: + rustfs-data: + gitea-data: +``` + +The volume-backed `/data` directory keeps the SQLite database and the Git repositories across container restarts, while LFS objects and attachments live in RustFS. + +## 2. Validate and start the deployment + +Resolve the Compose file before starting containers: + +```bash +docker compose config +``` + +Start the services: + +```bash +docker compose up -d +docker compose ps +``` + +Watch the Gitea log until every storage backend reports the Minio type: + +```bash +docker compose logs gitea | grep "Initialising" +``` + +The output should list `Attachment`, `Avatar`, `LFS`, and the remaining storage sections, each followed by a `Creating Minio storage at rustfs:9000:gitea-data` line. + +Open `http://localhost:3000` and create the administrator account, then open the RustFS Console at `http://localhost:9001` — the `gitea-data` bucket appears after the first storage operation. + +## 3. Push a Git LFS object + +Create a repository named `rustfs-demo` in the Gitea web UI, then push an LFS-tracked file from your workstation: + +```bash +mkdir lfs-demo && cd lfs-demo +git init +git config user.email you@example.com +git config user.name you +git lfs install +git lfs track "*.bin" +git add .gitattributes +dd if=/dev/urandom of=dataset.bin bs=1M count=8 +git add dataset.bin +git commit -m "add LFS dataset" +git remote add origin http://localhost:3000//rustfs-demo.git +git push origin main +``` + +`git push` uploads the LFS object through the Gitea LFS API, which writes it to RustFS. Clone the repository into a second directory and run `git lfs pull` — the downloaded `dataset.bin` must be byte-identical to the original: + +```bash +sha256sum dataset.bin +cd ../lfs-demo-clone && git lfs pull && sha256sum dataset.bin +``` + +Both checksums match because both clients read the object from RustFS. + +## 4. Attach a file to an issue + +Open the `rustfs-demo` repository, create an issue, and attach a small text file through the issue form. Gitea stores the upload as `attachments//` in the `gitea-data` bucket and serves downloads through `/attachments/`. + +## 5. Verify objects in RustFS + +List the bucket with the [`rc` client](https://github.com/rustfs/cli): + +```bash +docker compose exec rustfs /usr/bin/rc ls local/gitea-data/ -r +``` + +The output should include the LFS object under `lfs/` and the attachment under `attachments/`: + +```text +attachments/9/2/92fdd48d-531c-4cba-8b3f-4e2004a10fc7 +lfs/37/76/6ddfc07e803de58a69328db9a58a07cf7080ddde55c155a7531bc650a000 +``` + +The LFS object key is the SHA-256 content hash used by the Git LFS protocol. + +![Gitea LFS and attachment objects in the RustFS Console](./images/rustfs-gitea-objects.png) + +## 6. Stop or reset the deployment + +Stop the containers while keeping all data: + +```bash +docker compose down +``` + +The RustFS volume keeps the `gitea-data` bucket, so LFS objects and attachments survive a restart. To delete everything, including the objects in RustFS, add `--volumes`. + +## Troubleshooting + +### The Gitea install page appears instead of the login page + +The configuration file must exist at `/data/gitea/conf/app.ini` inside the container. If the mount path is wrong, Gitea starts with defaults and shows the installation wizard. Mount the file as shown in the Compose example and restart. + +### Push fails with an LFS or 403 error + +Confirm the credentials in `app.ini` match the RustFS credentials and that the `rustfs` hostname resolves inside the Compose network: + +```bash +docker compose logs gitea | grep -i minio +``` + +### Objects land in local storage instead of RustFS + +The `GITEA__storage__STORAGE_TYPE: minio` environment variable and the `[storage]` section of `app.ini` must agree. After changing either, restart Gitea and check the `Initialising` log lines again. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Gitea storage targets. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Gitea storage documentation](https://docs.gitea.com/administration/storage-configurations) to move packages, Actions artifacts, or individual storage sections to separate buckets. diff --git a/content/ja/developer/integration/devops/images/rustfs-gitea-objects.png b/content/ja/developer/integration/devops/images/rustfs-gitea-objects.png new file mode 100644 index 00000000..682f280c Binary files /dev/null and b/content/ja/developer/integration/devops/images/rustfs-gitea-objects.png differ diff --git a/content/ja/developer/integration/devops/index.md b/content/ja/developer/integration/devops/index.md index c77f6f04..caae08a1 100644 --- a/content/ja/developer/integration/devops/index.md +++ b/content/ja/developer/integration/devops/index.md @@ -8,6 +8,7 @@ S3 互換エンドポイントをサポートする DevOps プラットフォー ## プラットフォームとツール - [Elasticsearch](./elasticsearch.md) +- [Gitea](./gitea.md) - [Terraform](./terraform.md) アーティファクト、ステート、テレメトリデータは専用バケットに保存し、必要なバケット操作のみに権限が絞られた認証情報を使用してください。 diff --git a/content/ja/developer/integration/devops/meta.json b/content/ja/developer/integration/devops/meta.json index a032acbd..61ad9fb8 100644 --- a/content/ja/developer/integration/devops/meta.json +++ b/content/ja/developer/integration/devops/meta.json @@ -2,6 +2,7 @@ "title": "DevOps", "pages": [ "elasticsearch", + "gitea", "terraform" ] } diff --git a/content/ja/developer/integration/index.md b/content/ja/developer/integration/index.md index 075c44b3..e13a8768 100644 --- a/content/ja/developer/integration/index.md +++ b/content/ja/developer/integration/index.md @@ -8,11 +8,11 @@ description: "RustFS をリバースプロキシ、バックアップツール ## Integration categories - [Reverse Proxy](./reverse-proxy/index.md) は Nginx、Traefik、Caddy、HAProxy を扱います。 -- [Backup](./backup/index.md) は Restic を扱います。 +- [Backup](./backup/index.md) は Restic と Longhorn を扱います。 - [データ分析](./big-data/index.md) は Iceberg を扱います。 - [オブザーバビリティ](./observability/index.md) は OpenObserve を扱います。 - [その他](./others/index.md) はコミュニティ主導の Python 用 capo SDK を扱います。 - [コンテナレジストリ](./registry/index.md) は Harbor を扱います。 -- [DevOps](./devops/index.md) は Elasticsearch と Terraform を扱います。 +- [DevOps](./devops/index.md) は Elasticsearch、Gitea、Terraform を扱います。 各ガイドでは、連携先システムを設定する際に使用する RustFS のエンドポイントとアドレス指定の要件を示します。 \ No newline at end of file diff --git a/content/ja/developer/integration/observability/images/rustfs-thanos-blocks.png b/content/ja/developer/integration/observability/images/rustfs-thanos-blocks.png new file mode 100644 index 00000000..188cb9cf Binary files /dev/null and b/content/ja/developer/integration/observability/images/rustfs-thanos-blocks.png differ diff --git a/content/ja/developer/integration/observability/index.md b/content/ja/developer/integration/observability/index.md index 896cf131..f6485166 100644 --- a/content/ja/developer/integration/observability/index.md +++ b/content/ja/developer/integration/observability/index.md @@ -10,5 +10,6 @@ S3 互換エンドポイントをサポートするオブザーバビリティ - [OpenObserve](./openobserve.md) - [Loki](./loki.md) - [Tempo](./tempo.md) +- [Thanos](./thanos.md) テレメトリデータは専用バケットに保存し、必要なバケット操作のみに権限が絞られた認証情報を使用してください。 diff --git a/content/ja/developer/integration/observability/meta.json b/content/ja/developer/integration/observability/meta.json index f6d8d61c..16c15ac9 100644 --- a/content/ja/developer/integration/observability/meta.json +++ b/content/ja/developer/integration/observability/meta.json @@ -3,6 +3,7 @@ "pages": [ "openobserve", "loki", - "tempo" + "tempo", + "thanos" ] } diff --git a/content/ja/developer/integration/observability/thanos.md b/content/ja/developer/integration/observability/thanos.md new file mode 100644 index 00000000..856c31db --- /dev/null +++ b/content/ja/developer/integration/observability/thanos.md @@ -0,0 +1,311 @@ +--- +title: "Thanos" +description: "Run Thanos with RustFS as the S3 object storage backend for Prometheus blocks, deployed with Docker Compose." +--- + +This guide connects [Thanos](https://github.com/thanos-io/thanos) — the highly available Prometheus setup with long-term storage — to **RustFS** as its object store. You will run Prometheus with a Thanos sidecar that uploads TSDB blocks to RustFS, then query the historical data back through a Store Gateway and a Query frontend. The workflow was verified with `thanosio/thanos:v0.37.2`, `prom/prometheus:v2.53.1`, and `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with the Compose plugin. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Prom["Prometheus :9090"] -->|"blocks"| Sidecar["Thanos sidecar"] + Sidecar -->|"upload"| RustFS["RustFS :9000"] + Store["Store Gateway"] -->|"download"| RustFS + Query["Thanos Query"] -->|gRPC| Sidecar + Query -->|gRPC| Store +``` + +The sidecar watches the Prometheus TSDB directory and uploads every two-hour block to the `thanos-data` bucket in RustFS. The Store Gateway reads the same bucket and answers queries about historical blocks, so Query resolves both live data through the sidecar and old data through the Store Gateway. + +## 1. Create the project files + +Create a working directory: + +```bash +mkdir rustfs-thanos +cd rustfs-thanos +``` + +Create an environment file and replace both credential placeholders: + +```ini title=".env" +RUSTFS_ACCESS_KEY= +RUSTFS_SECRET_KEY= +``` + +Use dedicated credentials for the `thanos-data` bucket. Do not commit `.env` to source control. + +Create the Prometheus configuration with an external label — Thanos requires it to deduplicate blocks: + +```yaml title="prometheus.yml" +global: + scrape_interval: 5s + external_labels: + monitor: rustfs-demo + +scrape_configs: + - job_name: prometheus + static_configs: + - targets: ["localhost:9090"] + - job_name: rustfs + metrics_path: /metrics + static_configs: + - targets: ["rustfs:9000"] +``` + +Create the Thanos object store configuration: + +```yaml title="bucket.yml" +type: S3 +config: + bucket: thanos-data + endpoint: rustfs:9000 + access_key: ${RUSTFS_ACCESS_KEY} + secret_key: ${RUSTFS_SECRET_KEY} + insecure: true +``` + +Thanos does not interpolate `.env` files itself. Before starting the stack, replace the placeholders with the same values you set in `.env`: + +```bash +sed -i.bak "s|\${RUSTFS_ACCESS_KEY}|$(grep RUSTFS_ACCESS_KEY .env | cut -d= -f2)|;s|\${RUSTFS_SECRET_KEY}|$(grep RUSTFS_SECRET_KEY .env | cut -d= -f2)|" bucket.yml +``` + +Create the Compose file: + +```yaml title="compose.yaml" +services: + rustfs: + image: rustfs/rustfs-x86-musl:v2.3.1 + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + RUSTFS_VOLUMES: /data + RUSTFS_ADDRESS: ":9000" + RUSTFS_CONSOLE_ADDRESS: ":9001" + RUSTFS_CONSOLE_ENABLE: "true" + volumes: + - rustfs-data:/data + ports: + - "9000:9000" + - "9001:9001" + healthcheck: + test: ["CMD", "curl", "-sf", "http://127.0.0.1:9000/health"] + interval: 10s + timeout: 5s + retries: 6 + start_period: 10s + networks: + - thanos + + create-bucket: + image: rustfs/rc:latest + depends_on: + rustfs: + condition: service_healthy + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + entrypoint: + - /bin/sh + - -c + - | + /usr/bin/rc alias set rustfs http://rustfs:9000 "$${RUSTFS_ACCESS_KEY}" "$${RUSTFS_SECRET_KEY}" + /usr/bin/rc mb --ignore-existing rustfs/thanos-data + networks: + - thanos + + prometheus: + image: prom/prometheus:v2.53.1 + command: + - --config.file=/etc/prometheus/prometheus.yml + - --storage.tsdb.path=/prometheus + - --storage.tsdb.min-block-duration=2h + - --storage.tsdb.max-block-duration=2h + - --web.enable-lifecycle + volumes: + - ./prometheus.yml:/etc/prometheus/prometheus.yml:ro + - prom-data:/prometheus + ports: + - "9090:9090" + networks: + - thanos + + sidecar: + image: thanosio/thanos:v0.37.2 + command: + - sidecar + - --tsdb.path=/prometheus + - --prometheus.url=http://prometheus:9090 + - --objstore.config-file=/etc/thanos/bucket.yml + volumes: + - ./bucket.yml:/etc/thanos/bucket.yml:ro + - prom-data:/prometheus + depends_on: + create-bucket: + condition: service_completed_successfully + networks: + - thanos + + store: + image: thanosio/thanos:v0.37.2 + command: + - store + - --objstore.config-file=/etc/thanos/bucket.yml + - --data-dir=/data + volumes: + - ./bucket.yml:/etc/thanos/bucket.yml:ro + - store-data:/data + depends_on: + create-bucket: + condition: service_completed_successfully + networks: + - thanos + + query: + image: thanosio/thanos:v0.37.2 + command: + - query + - --http-address=0.0.0.0:9090 + - --store=sidecar:10901 + - --store=store:10901 + ports: + - "9091:9090" + depends_on: + - sidecar + - store + networks: + - thanos + +networks: + thanos: + +volumes: + rustfs-data: + prom-data: + store-data: +``` + +The `--storage.tsdb.min-block-duration` and `--storage.tsdb.max-block-duration` flags disable Prometheus compaction. The sidecar refuses to ship blocks from a compacting TSDB because the local blocks would no longer match the uploaded ones. + +## 2. Validate and start the deployment + +Resolve the Compose file before starting containers: + +```bash +docker compose config +``` + +Start the stack and wait until the sidecar reports itself ready: + +```bash +docker compose up -d +docker compose logs sidecar | grep -m1 "status=ready" +``` + +Check that the sidecar picked up the Prometheus external labels: + +```bash +docker compose logs sidecar | grep "external labels" +``` + +The Thanos Query UI answers on `http://localhost:9091`, and the RustFS Console runs at `http://localhost:9001`. + +## 3. Upload a block to RustFS + +The sidecar uploads a block when Prometheus compacts one, which happens at a two-hour block boundary. To produce a block immediately, snapshot the TSDB through the admin API — with compaction disabled, the sidecar ships the head-block snapshot directly: + +```bash +curl -s -XPOST http://localhost:9090/api/v1/admin/tsdb/snapshot | head -c 200 +``` + +Wait for the upload, then check the shipper state inside Prometheus: + +```bash +sleep 60 +docker compose exec prometheus cat /prometheus/thanos.shipper.json +``` + +The `uploaded` list should contain a block ID: + +```json +{ + "version": 1, + "uploaded": [ + "01M31EPTZC5E0SETZTP0SPFY79" + ] +} +``` + +## 4. Query historical data from RustFS + +The Store Gateway periodically syncs the bucket. Confirm it downloaded the uploaded block: + +```bash +docker compose logs store | grep "loaded new block" +``` + +Query a series through the Query frontend over the block's time range: + +```bash +START=$(date -u -d '2 hours ago' +%s) +END=$(date -u +%s) +curl -s "http://localhost:9091/api/v1/query_range?query=up%7Bjob%3D%22prometheus%22%7D&start=$START&end=$END&step=30" | head -c 300 +``` + +The Store Gateway serves the response from the blocks it downloaded from RustFS, while the sidecar answers for the live head — both paths resolve through the same Query endpoint. + +## 5. Verify objects in RustFS + +List the bucket: + +```bash +docker compose exec rustfs /usr/bin/rc ls local/thanos-data/ -r +``` + +Each block is stored as three objects — the chunk files, the index, and `meta.json`: + +```text +01M31EPTZC5E0SETZTP0SPFY79/chunks/000001 +01M31EPTZC5E0SETZTP0SPFY79/index +01M31EPTZC5E0SETZTP0SPFY79/meta.json +``` + +![Thanos blocks stored in the RustFS Console](./images/rustfs-thanos-blocks.png) + +## 6. Stop or reset the deployment + +Stop the containers while keeping all data: + +```bash +docker compose down +``` + +The RustFS volume keeps the uploaded blocks, so the Store Gateway serves historical queries again after a restart. To delete everything, including the blocks in RustFS, add `--volumes`. + +## Troubleshooting + +### The sidecar logs `Compaction needs to be disabled` + +Prometheus must run with `--storage.tsdb.min-block-duration` equal to `--storage.tsdb.max-block-duration` — set both to `2h` as shown in the Compose file. Otherwise the sidecar cannot guarantee that local blocks stay unchanged and refuses to upload. + +### `The specified bucket does not exist` + +Thanos does not create buckets. Check that the `create-bucket` service completed successfully: + +```bash +docker compose logs create-bucket +``` + +### Queries return no historical data + +Confirm that the Store Gateway has loaded at least one block (`docker compose logs store | grep "loaded new block"`) and that your query time range falls inside the uploaded block's window — check the block's `meta.json` in the RustFS Console for `minTime` and `maxTime`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Thanos components. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Thanos documentation](https://thanos.io/tip/thanos/getting-started.md) to add Compactor, Ruler, or Receive for a production topology. diff --git a/content/zh/developer/integration/backup/images/rustfs-longhorn-backups.png b/content/zh/developer/integration/backup/images/rustfs-longhorn-backups.png new file mode 100644 index 00000000..cc577a01 Binary files /dev/null and b/content/zh/developer/integration/backup/images/rustfs-longhorn-backups.png differ diff --git a/content/zh/developer/integration/backup/index.md b/content/zh/developer/integration/backup/index.md index 14e98539..f56e073d 100644 --- a/content/zh/developer/integration/backup/index.md +++ b/content/zh/developer/integration/backup/index.md @@ -8,5 +8,6 @@ description: "通过兼容 S3 的对象存储接口将备份工具连接到 Rust ## 系统 - [Restic](./restic.md) +- [Longhorn](./longhorn.md) 请将备份作业放在专用的存储桶和前缀中,并使用仅限所需存储桶操作的凭据。 \ No newline at end of file diff --git a/content/zh/developer/integration/backup/longhorn.md b/content/zh/developer/integration/backup/longhorn.md new file mode 100644 index 00000000..272a73c6 --- /dev/null +++ b/content/zh/developer/integration/backup/longhorn.md @@ -0,0 +1,293 @@ +--- +title: "Longhorn" +description: "配置 Longhorn 通过其 S3 备份目标把 Kubernetes 卷备份存储到 RustFS。" +--- + +本指南将 Kubernetes 的分布式块存储系统 [Longhorn](https://github.com/longhorn/longhorn) 连接到 **RustFS** 作为其 S3 备份目标。你将配置备份目标,备份一个包含数据的卷,删除该卷后从 RustFS 恢复,并验证数据。整个流程使用 Longhorn 1.9.0(k3s,Kubernetes 1.30)和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要一个已安装 Longhorn 的 Kubernetes 集群,并具备 `kubectl` 访问权限。本指南用于集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + App["Workload pod"] -->|"writes"| Vol["Longhorn volume"] + Vol -->|"snapshot"| Backup["Backup engine"] + Backup -->|"blocks + config"| RustFS["RustFS :9000"] +``` + +Longhorn 把备份以内容寻址块和一个小的卷配置文件的形式,存到目标存储桶的 `backupstore/volumes/` 前缀下。任何能访问 RustFS 的节点都可以把这些对象读回一个新卷来完成恢复。 + +## 1. 创建备份存储桶 + +使用 [`rc` 客户端](https://github.com/rustfs/cli)创建专用存储桶,替换为你自己的端点和凭证: + +```bash +rc alias set rustfs http://:9000 +rc mb rustfs/longhorn-backups +``` + +Longhorn 不会创建存储桶,因此这一步必须在第一次备份之前完成。 + +## 2. 配置备份目标 + +把 RustFS 凭证保存为 `longhorn-system` 命名空间中的 Secret。`AWS_ENDPOINTS` 必须是每个节点都能访问的端点——请使用节点 IP 或内部负载均衡地址,而不是你工作站上的端口转发: + +```bash +kubectl -n longhorn-system create secret generic rustfs-s3-secret \ + --from-literal=AWS_ACCESS_KEY_ID= \ + --from-literal=AWS_SECRET_ACCESS_KEY= \ + --from-literal=AWS_ENDPOINTS=http://:9000 +``` + +Longhorn 1.9 通过 `BackupTarget` 资源管理备份目标。用 RustFS 存储桶修补 `default` 目标: + +```bash +kubectl -n longhorn-system patch backupTarget default --type merge -p ' +spec: + backupTargetURL: s3://longhorn-backups@us-east-1/ + credentialSecret: rustfs-s3-secret + pollInterval: 5m' +``` + +`@us-east-1` 段是 S3 URL 格式中的区域标注,不需要与实际部署区域一致。 + +等待目标变为可用——它确认 Longhorn 已通过该 Secret 访问到存储桶: + +```bash +kubectl -n longhorn-system get backupTarget default +``` + +```text +NAME URL CREDENTIAL AVAILABLE LASTSYNCEDAT +default s3://longhorn-backups@us-east-1/ rustfs-s3-secret true 2026-09-21T08:36:55Z +``` + +## 3. 写入数据并创建备份 + +创建一个包含数据的测试卷。以下清单创建一个 1 GiB 的 PVC 和一个写入标记文件的 Pod: + +```yaml title="demo.yaml" +apiVersion: v1 +kind: PersistentVolumeClaim +metadata: + name: demo-vol +spec: + accessModes: [ReadWriteOnce] + storageClassName: longhorn + resources: + requests: + storage: 1Gi +--- +apiVersion: v1 +kind: Pod +metadata: + name: demo-app +spec: + volumes: + - name: data + persistentVolumeClaim: + claimName: demo-vol + containers: + - name: app + image: busybox:1.36 + command: ["sh", "-c", "echo 'longhorn rustfs demo' > /data/hello.txt && sleep 3600"] + volumeMounts: + - name: data + mountPath: /data +``` + +应用并等待 Pod 运行: + +```bash +kubectl apply -f demo.yaml +kubectl get pod demo-app +``` + +创建快照并备份。你可以在 Longhorn UI 中操作(Volume → Snapshot → Backup),也可以用声明式方式: + +```bash +VOLUME=$(kubectl get pvc demo-vol -o jsonpath='{.spec.volumeName}') +kubectl -n longhorn-system apply -f - <" +EOF +``` + +恢复在卷首次 attach 时执行。通过静态 PV 创建绑定到恢复卷的 PVC,然后挂载它: + +```bash +kubectl apply -f - < --type merge -p ' +spec: + disks: + default: + path: /var/lib/longhorn + allowScheduling: true' +``` + +### `failed to create backup ... missing input parameter` + +备份在默认备份目标配置完成之前发起。先完成第 2 步,确认 `AVAILABLE` 为 `true`,再重新创建备份。 + +### 备份目标始终不可用 + +Secret 必须在目标同步之前存在,并且 `AWS_ENDPOINTS` 必须从节点本身可达。检查 `longhorn-manager` 日志中的 S3 错误: + +```bash +kubectl -n longhorn-system logs -l app=longhorn-manager | grep -i s3 | tail +``` + +## 后续步骤 + +- 在采用更多备份目标之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Longhorn 备份文档](https://longhorn.io/docs/1.9.0/backups-and-restore/)配置周期性备份作业和计划快照。 diff --git a/content/zh/developer/integration/backup/meta.json b/content/zh/developer/integration/backup/meta.json index 51649688..4511320b 100644 --- a/content/zh/developer/integration/backup/meta.json +++ b/content/zh/developer/integration/backup/meta.json @@ -1,6 +1,7 @@ { "title": "备份", "pages": [ - "restic" + "restic", + "longhorn" ] -} \ No newline at end of file +} diff --git a/content/zh/developer/integration/big-data/images/rustfs-mlflow-artifacts.png b/content/zh/developer/integration/big-data/images/rustfs-mlflow-artifacts.png new file mode 100644 index 00000000..0d463bea Binary files /dev/null and b/content/zh/developer/integration/big-data/images/rustfs-mlflow-artifacts.png differ diff --git a/content/zh/developer/integration/big-data/index.md b/content/zh/developer/integration/big-data/index.md index c2b38258..73189d80 100644 --- a/content/zh/developer/integration/big-data/index.md +++ b/content/zh/developer/integration/big-data/index.md @@ -10,6 +10,7 @@ description: "通过 S3 兼容的对象存储接口将数据分析系统连接 - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) - [Milvus](./milvus.md) +- [MLflow](./mlflow.md) - [DuckDB](./duckdb.md) - [InfluxDB](./influxdb.md) - [Spark](./spark.md) diff --git a/content/zh/developer/integration/big-data/meta.json b/content/zh/developer/integration/big-data/meta.json index 448f23d8..a951cb56 100644 --- a/content/zh/developer/integration/big-data/meta.json +++ b/content/zh/developer/integration/big-data/meta.json @@ -4,6 +4,7 @@ "iceberg", "pyiceberg", "milvus", + "mlflow", "duckdb", "influxdb", "spark", diff --git a/content/zh/developer/integration/big-data/mlflow.md b/content/zh/developer/integration/big-data/mlflow.md new file mode 100644 index 00000000..ef6d644f --- /dev/null +++ b/content/zh/developer/integration/big-data/mlflow.md @@ -0,0 +1,244 @@ +--- +title: "MLflow" +description: "使用 Docker Compose 部署 MLflow,以 RustFS 作为实验跟踪的 S3 制品存储。" +--- + +本指南将实验跟踪与模型注册平台 [MLflow](https://github.com/mlflow/mlflow) 连接到 **RustFS**,作为其 S3 制品存储。你将使用 Docker Compose 启动 MLflow 跟踪服务器,记录一次训练运行的参数、指标和制品,然后验证制品已存储在 RustFS 中。整个流程使用 `ghcr.io/mlflow/mlflow:v2.22.1` 和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装带有 Compose 插件的 Docker。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Client["Training script"] -->|"runs + metrics"| Server["MLflow server :5000"] + Client -->|"artifacts"| RustFS["RustFS :9000"] + Server -->|"metadata"| DB["SQLite"] +``` + +跟踪服务器将实验和运行元数据保存在 SQLite 中,而制品——模型文件、图表、报告——通过 `s3://` 制品根路径直接存入 RustFS。客户端也需要同样的 RustFS 凭证,因为制品由客户端自己上传,借助 boto3 和 `MLFLOW_S3_ENDPOINT_URL` 设置完成。 + +## 1. 创建项目文件 + +创建工作目录: + +```bash +mkdir rustfs-mlflow +cd rustfs-mlflow +``` + +创建环境文件,并替换两个凭证占位符: + +```ini title=".env" +RUSTFS_ACCESS_KEY= +RUSTFS_SECRET_KEY= +MLFLOW_BUCKET=my-bucket +``` + +请为制品存储桶使用专用凭证。不要将 `.env` 提交到版本控制。 + +创建 Compose 文件: + +```yaml title="compose.yaml" +services: + rustfs: + image: rustfs/rustfs-x86-musl:v2.3.1 + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + RUSTFS_VOLUMES: /data + RUSTFS_ADDRESS: ":9000" + RUSTFS_CONSOLE_ADDRESS: ":9001" + RUSTFS_CONSOLE_ENABLE: "true" + volumes: + - rustfs-data:/data + ports: + - "9000:9000" + - "9001:9001" + healthcheck: + test: ["CMD", "curl", "-sf", "http://127.0.0.1:9000/health"] + interval: 10s + timeout: 5s + retries: 6 + start_period: 10s + networks: + - mlflow + + create-bucket: + image: rustfs/rc:latest + depends_on: + rustfs: + condition: service_healthy + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + MLFLOW_BUCKET: ${MLFLOW_BUCKET} + entrypoint: + - /bin/sh + - -c + - | + /usr/bin/rc alias set rustfs http://rustfs:9000 "$${RUSTFS_ACCESS_KEY}" "$${RUSTFS_SECRET_KEY}" + /usr/bin/rc mb --ignore-existing rustfs/$${MLFLOW_BUCKET} + networks: + - mlflow + + mlflow: + image: ghcr.io/mlflow/mlflow:v2.22.1 + command: + - server + - --backend-store-uri + - sqlite:////mlflow/mlflow.db + - --default-artifact-root + - s3://${MLFLOW_BUCKET}/mlflow-artifacts + - --host + - 0.0.0.0 + - --port + - "5000" + environment: + AWS_ACCESS_KEY_ID: ${RUSTFS_ACCESS_KEY} + AWS_SECRET_ACCESS_KEY: ${RUSTFS_SECRET_KEY} + MLFLOW_S3_ENDPOINT_URL: http://rustfs:9000 + AWS_DEFAULT_REGION: us-east-1 + depends_on: + create-bucket: + condition: service_completed_successfully + ports: + - "5000:5000" + volumes: + - mlflow-db:/mlflow + networks: + - mlflow + +networks: + mlflow: + +volumes: + rustfs-data: + mlflow-db: +``` + +`create-bucket` 服务必须先于服务器运行,因为 MLflow 不会创建存储桶。`MLFLOW_S3_ENDPOINT_URL` 将服务器的 boto3 客户端指向 RustFS,并以 path-style 方式寻址。SQLite 数据库保存在卷中,实验元数据因此可以在重启后保留;生产部署应改用受管数据库后端。 + +## 2. 校验并启动部署 + +启动容器前先解析 Compose 文件: + +```bash +docker compose config +``` + +启动服务并等待跟踪服务器就绪: + +```bash +docker compose up -d +docker compose ps +``` + +`create-bucket` 服务应以退出码 `0` 结束,MLflow UI 应在 `http://localhost:5000` 上响应: + +```bash +curl -sf http://localhost:5000/ >/dev/null && echo ready +``` + +打开 `http://localhost:9001` 的 RustFS 控制台,在下一步中观察制品落入 `my-bucket`。 + +## 3. 记录一次训练运行 + +创建客户端脚本——它在 MLflow 镜像内运行,镜像自带全部依赖: + +```bash title="train_demo.py" {12} +import mlflow + +mlflow.set_tracking_uri("http://localhost:5000") +mlflow.set_experiment("rustfs-demo") + +with mlflow.start_run(run_name="rustfs-verify") as run: + mlflow.log_params({"model": "demo-regressor", "alpha": 0.5}) + for step in range(3): + mlflow.log_metric("rmse", 0.9 - step * 0.2, step=step) + with open("model-summary.txt", "w") as f: + f.write("demo model trained against RustFS artifact store\n") + mlflow.log_artifact("model-summary.txt", artifact_path="reports") + print("run_id:", run.info.run_id) + print("artifact_uri:", run.info.artifact_uri) +``` + +把脚本复制到运行中的容器里,并使用服务环境执行: + +```bash +docker compose cp train_demo.py mlflow:/tmp/train_demo.py +docker compose exec -w /tmp mlflow python train_demo.py +``` + +`artifact_uri` 会输出为 `s3://my-bucket/mlflow-artifacts///artifacts`——制品由客户端直接上传到 RustFS。 + +## 4. 在 RustFS 中验证制品 + +通过跟踪服务器读回制品,然后列出 RustFS 中的同一对象。将以下内容保存为 `verify.py`,把 `` 替换为第 3 步输出的标识符,并以相同方式运行: + +```bash title="verify.py" {5} +import mlflow + +mlflow.set_tracking_uri("http://localhost:5000") +client = mlflow.MlflowClient() +print([a.path for a in client.list_artifacts("", "reports")]) +path = client.download_artifacts("", "reports/model-summary.txt") +print(open(path).read()) +``` + +下载操作会通过跟踪服务器从 RustFS 读取对象。然后确认存储桶中的对象: + +```bash +docker compose exec rustfs /usr/bin/rc ls local/my-bucket/mlflow-artifacts/ -r +``` + +输出应包含该制品对象: + +```text +mlflow-artifacts/1/a1aece9243504f2680a52fba0c32765f/artifacts/reports/model-summary.txt +``` + +![RustFS 控制台中存储的 MLflow 制品](./images/rustfs-mlflow-artifacts.png) + +运行、参数和指标在服务器重启后依然保留,因为它们存储在 SQLite 中,而制品保存在 RustFS 里: + +```bash +docker compose restart mlflow +docker compose exec -w /tmp mlflow python verify.py +``` + +重启后的服务器上 `list_artifacts` 调用依然成功,读取的是 RustFS 中相同的对象。 + +## 5. 停止或重置部署 + +停止容器并保留所有数据: + +```bash +docker compose down +``` + +RustFS 卷会保留制品对象,MLflow 卷会保留元数据数据库。若要删除包括 RustFS 中制品在内的所有数据,请追加 `--volumes`。 + +## 故障排查 + +### 记录制品时出现 `ModuleNotFoundError: No module named 'boto3'` + +执行 `log_artifact` 的客户端需要 boto3,因为制品是直接上传到 S3 的。在运行训练脚本的环境中安装它,或按第 3 步所示在 MLflow 镜像内运行脚本。 + +### 制品上传时出现 `AccessDenied` 或连接错误 + +确认客户端环境中设置了 `MLFLOW_S3_ENDPOINT_URL`——缺少它时,boto3 会把请求发到真实的 AWS S3。端点必须从运行训练脚本的机器可达;在 Compose 网络外使用 `http://localhost:9000`,网络内使用 `http://rustfs:9000`。 + +### 服务器启动失败并报存储桶错误 + +制品存储桶必须在服务器启动前存在。检查 `create-bucket` 服务的日志: + +```bash +docker compose logs create-bucket +``` + +## 后续步骤 + +- 在采用更多 MLflow 操作之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [MLflow 文档](https://mlflow.org/docs/latest/)添加模型注册表,或将元数据存储迁移到受管数据库。 diff --git a/content/zh/developer/integration/devops/gitea.md b/content/zh/developer/integration/devops/gitea.md new file mode 100644 index 00000000..928c76f8 --- /dev/null +++ b/content/zh/developer/integration/devops/gitea.md @@ -0,0 +1,236 @@ +--- +title: "Gitea" +description: "使用 Docker Compose 部署 Gitea,以 RustFS 作为 LFS 对象和附件的 S3 存储后端。" +--- + +本指南将自托管 Git 服务 [Gitea](https://github.com/go-gitea/gitea) 通过其 `minio` 存储类型连接到 **RustFS**。你将使用 Docker Compose 启动 Gitea,创建仓库并推送 Git LFS 对象,在 issue 中上传附件,然后验证 LFS 对象和附件都存储在 RustFS 中。整个流程使用 `gitea/gitea:1.24.4` 和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装带有 Compose 插件的 Docker,工作站上需要有 `git` 和 `git-lfs` 客户端。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Dev["Git + LFS client"] -->|"git push / git-lfs"| Gitea["Gitea :3000"] + Gitea -->|"LFS + attachments"| RustFS["RustFS :9000"] +``` + +Gitea 将 Git 仓库本身保存在本地磁盘上,而 `minio` 存储类型会把大文件——LFS 对象、issue 附件、头像、仓库归档、软件包和 Actions 制品——路由到 RustFS 的 `gitea-data` 存储桶。如果桶不存在,Gitea 会在启动时自动创建。 + +## 1. 创建项目文件 + +创建工作目录: + +```bash +mkdir rustfs-gitea +cd rustfs-gitea +``` + +创建环境文件,并替换两个凭证占位符: + +```ini title=".env" +RUSTFS_ACCESS_KEY= +RUSTFS_SECRET_KEY= +``` + +请为 `gitea-data` 存储桶使用专用凭证。不要将 `.env` 提交到版本控制。 + +创建 Gitea 配置文件,并将两个凭证占位符替换为与 `.env` 相同的值——全局 `minio` 存储类型适用于 LFS、附件、头像、仓库归档、软件包和 Actions 制品: + +```ini title="app.ini" +APP_NAME = RustFS Gitea +RUN_MODE = prod +WORK_PATH = /data/gitea + +[server] +DOMAIN = localhost +ROOT_URL = http://localhost:3000/ +HTTP_PORT = 3000 +LFS_START_SERVER = true + +[database] +DB_TYPE = sqlite3 +PATH = /data/gitea/gitea.db + +[storage] +STORAGE_TYPE = minio +MINIO_ENDPOINT = rustfs:9000 +MINIO_ACCESS_KEY_ID = +MINIO_SECRET_ACCESS_KEY = +MINIO_BUCKET = gitea-data +MINIO_LOCATION = us-east-1 +MINIO_USE_SSL = false + +[log] +MODE = console +LEVEL = info + +[security] +INSTALL_LOCK = true +SECRET_KEY = change-me-to-a-random-string +``` + +`LFS_START_SERVER` 用于启用 Git LFS HTTP API。`MINIO_ENDPOINT` 使用 Compose 网络内的主机名 `rustfs`;该端点默认以 path-style 方式访问。 + +创建 Compose 文件: + +```yaml title="compose.yaml" +services: + rustfs: + image: rustfs/rustfs-x86-musl:v2.3.1 + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + RUSTFS_VOLUMES: /data + RUSTFS_ADDRESS: ":9000" + RUSTFS_CONSOLE_ADDRESS: ":9001" + RUSTFS_CONSOLE_ENABLE: "true" + volumes: + - rustfs-data:/data + ports: + - "9000:9000" + - "9001:9001" + healthcheck: + test: ["CMD", "curl", "-sf", "http://127.0.0.1:9000/health"] + interval: 10s + timeout: 5s + retries: 6 + start_period: 10s + networks: + - gitea + + gitea: + image: gitea/gitea:1.24.4 + depends_on: + rustfs: + condition: service_healthy + environment: + USER_UID: "1000" + USER_GID: "1000" + volumes: + - gitea-data:/data + - ./app.ini:/data/gitea/conf/app.ini:ro + ports: + - "3000:3000" + networks: + - gitea + +networks: + gitea: + +volumes: + rustfs-data: + gitea-data: +``` + +由卷承载的 `/data` 目录让 SQLite 数据库和 Git 仓库在容器重启后得以保留,而 LFS 对象和附件则存放在 RustFS 中。 + +## 2. 校验并启动部署 + +启动容器前先解析 Compose 文件: + +```bash +docker compose config +``` + +启动服务: + +```bash +docker compose up -d +docker compose ps +``` + +观察 Gitea 日志,直到每个存储后端都报告为 Minio 类型: + +```bash +docker compose logs gitea | grep "Initialising" +``` + +输出应列出 `Attachment`、`Avatar`、`LFS` 等存储段,每段后面跟着一行 `Creating Minio storage at rustfs:9000:gitea-data`。 + +打开 `http://localhost:3000` 创建管理员账号,然后打开 `http://localhost:9001` 的 RustFS 控制台——第一次存储操作后会出现 `gitea-data` 存储桶。 + +## 3. 推送 Git LFS 对象 + +在 Gitea 网页界面创建名为 `rustfs-demo` 的仓库,然后从工作站推送一个 LFS 跟踪的文件: + +```bash +mkdir lfs-demo && cd lfs-demo +git init +git config user.email you@example.com +git config user.name you +git lfs install +git lfs track "*.bin" +git add .gitattributes +dd if=/dev/urandom of=dataset.bin bs=1M count=8 +git add dataset.bin +git commit -m "add LFS dataset" +git remote add origin http://localhost:3000//rustfs-demo.git +git push origin main +``` + +`git push` 会通过 Gitea 的 LFS API 上传 LFS 对象,Gitea 将其写入 RustFS。把仓库克隆到另一个目录并执行 `git lfs pull`——下载的 `dataset.bin` 必须与原始文件逐字节一致: + +```bash +sha256sum dataset.bin +cd ../lfs-demo-clone && git lfs pull && sha256sum dataset.bin +``` + +两个校验和一致,因为两个客户端都从 RustFS 读取对象。 + +## 4. 在 issue 中上传附件 + +打开 `rustfs-demo` 仓库,创建一个 issue,并通过 issue 表单上传一个小的文本文件。Gitea 会将上传内容以 `attachments/<前缀>/` 的形式存入 `gitea-data` 存储桶,并通过 `/attachments/` 提供下载。 + +## 5. 在 RustFS 中验证对象 + +使用 [`rc` 客户端](https://github.com/rustfs/cli)列出存储桶: + +```bash +docker compose exec rustfs /usr/bin/rc ls local/gitea-data/ -r +``` + +输出应包含 `lfs/` 下的 LFS 对象和 `attachments/` 下的附件: + +```text +attachments/9/2/92fdd48d-531c-4cba-8b3f-4e2004a10fc7 +lfs/37/76/6ddfc07e803de58a69328db9a58a07cf7080ddde55c155a7531bc650a000 +``` + +LFS 对象的键是 Git LFS 协议使用的 SHA-256 内容哈希。 + +![RustFS 控制台中的 Gitea LFS 对象与附件](./images/rustfs-gitea-objects.png) + +## 6. 停止或重置部署 + +停止容器并保留所有数据: + +```bash +docker compose down +``` + +RustFS 卷会保留 `gitea-data` 存储桶,LFS 对象和附件在重启后依然可用。若要删除包括 RustFS 中对象在内的所有数据,请追加 `--volumes`。 + +## 故障排查 + +### 显示的是 Gitea 安装页面而不是登录页面 + +配置文件必须存在于容器内的 `/data/gitea/conf/app.ini`。如果挂载路径不对,Gitea 会以默认配置启动并显示安装向导。按 Compose 示例挂载该文件并重启。 + +### 推送失败并出现 LFS 或 403 错误 + +确认 `app.ini` 中的凭证与 RustFS 的凭证一致,并且 `rustfs` 主机名能在 Compose 网络内解析: + +```bash +docker compose logs gitea | grep -i minio +``` + +### 对象写进了本地存储而不是 RustFS + +`GITEA__storage__STORAGE_TYPE: minio` 环境变量与 `app.ini` 的 `[storage]` 段必须一致。修改任意一处后,重启 Gitea 并重新检查 `Initialising` 日志行。 + +## 后续步骤 + +- 在采用更多 Gitea 存储目标之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Gitea 存储文档](https://docs.gitea.com/administration/storage-configurations)将软件包、Actions 制品或单个存储段迁移到独立的存储桶。 diff --git a/content/zh/developer/integration/devops/images/rustfs-gitea-objects.png b/content/zh/developer/integration/devops/images/rustfs-gitea-objects.png new file mode 100644 index 00000000..c0aacff5 Binary files /dev/null and b/content/zh/developer/integration/devops/images/rustfs-gitea-objects.png differ diff --git a/content/zh/developer/integration/devops/index.md b/content/zh/developer/integration/devops/index.md index 396b1f28..dd023482 100644 --- a/content/zh/developer/integration/devops/index.md +++ b/content/zh/developer/integration/devops/index.md @@ -8,6 +8,7 @@ description: "通过 S3 兼容的对象存储接口,将 DevOps 平台与基础 ## 平台与工具 - [Elasticsearch](./elasticsearch.md) +- [Gitea](./gitea.md) - [Terraform](./terraform.md) 请使用专用的存储桶保存制品、状态与遥测数据,并为凭证仅授予所需桶操作的权限。 diff --git a/content/zh/developer/integration/devops/meta.json b/content/zh/developer/integration/devops/meta.json index a032acbd..61ad9fb8 100644 --- a/content/zh/developer/integration/devops/meta.json +++ b/content/zh/developer/integration/devops/meta.json @@ -2,6 +2,7 @@ "title": "DevOps", "pages": [ "elasticsearch", + "gitea", "terraform" ] } diff --git a/content/zh/developer/integration/index.md b/content/zh/developer/integration/index.md index d5c4a62f..82b9101f 100644 --- a/content/zh/developer/integration/index.md +++ b/content/zh/developer/integration/index.md @@ -8,11 +8,11 @@ description: "将 RustFS 与反向代理、备份工具、数据分析系统、 ## 集成类别 - [反向代理](./reverse-proxy/index.md)涵盖 Nginx、Traefik、Caddy 和 HAProxy。 -- [备份](./backup/index.md)涵盖 Restic。 +- [备份](./backup/index.md)涵盖 Restic 和 Longhorn。 - [数据分析](./big-data/index.md)涵盖 Iceberg。 - [可观测性](./observability/index.md)涵盖 OpenObserve。 - [其他](./others/index.md)涵盖社区驱动的 Python capo SDK。 - [镜像仓库](./registry/index.md)涵盖 Harbor。 -- [DevOps](./devops/index.md)涵盖 Elasticsearch 和 Terraform。 +- [DevOps](./devops/index.md)涵盖 Elasticsearch、Gitea 和 Terraform。 每篇指南都会说明配置集成系统时需要使用的 RustFS 端点和寻址要求。 \ No newline at end of file diff --git a/content/zh/developer/integration/observability/images/rustfs-thanos-blocks.png b/content/zh/developer/integration/observability/images/rustfs-thanos-blocks.png new file mode 100644 index 00000000..9e0aca22 Binary files /dev/null and b/content/zh/developer/integration/observability/images/rustfs-thanos-blocks.png differ diff --git a/content/zh/developer/integration/observability/index.md b/content/zh/developer/integration/observability/index.md index 9234d46d..5d85e502 100644 --- a/content/zh/developer/integration/observability/index.md +++ b/content/zh/developer/integration/observability/index.md @@ -10,5 +10,6 @@ description: "通过 S3 兼容对象存储接口,将可观测性平台连接 - [OpenObserve](./openobserve.md) - [Loki](./loki.md) - [Tempo](./tempo.md) +- [Thanos](./thanos.md) 请使用专用的存储桶保存遥测数据,并为凭证仅授予所需桶操作的权限。 diff --git a/content/zh/developer/integration/observability/meta.json b/content/zh/developer/integration/observability/meta.json index 7efedc39..4db188bf 100644 --- a/content/zh/developer/integration/observability/meta.json +++ b/content/zh/developer/integration/observability/meta.json @@ -3,6 +3,7 @@ "pages": [ "openobserve", "loki", - "tempo" + "tempo", + "thanos" ] } diff --git a/content/zh/developer/integration/observability/thanos.md b/content/zh/developer/integration/observability/thanos.md new file mode 100644 index 00000000..476b8efb --- /dev/null +++ b/content/zh/developer/integration/observability/thanos.md @@ -0,0 +1,311 @@ +--- +title: "Thanos" +description: "使用 Docker Compose 部署 Thanos,以 RustFS 作为 Prometheus 块的 S3 对象存储后端。" +--- + +本指南将 [Thanos](https://github.com/thanos-io/thanos)——具备长期存储的高可用 Prometheus 方案——连接到 **RustFS** 作为其对象存储。你将运行带 Thanos sidecar 的 Prometheus,sidecar 会把 TSDB 块上传到 RustFS,然后通过 Store Gateway 和 Query 前端把历史数据查询回来。整个流程使用 `thanosio/thanos:v0.37.2`、`prom/prometheus:v2.53.1` 和 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装带有 Compose 插件的 Docker。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Prom["Prometheus :9090"] -->|"blocks"| Sidecar["Thanos sidecar"] + Sidecar -->|"upload"| RustFS["RustFS :9000"] + Store["Store Gateway"] -->|"download"| RustFS + Query["Thanos Query"] -->|gRPC| Sidecar + Query -->|gRPC| Store +``` + +sidecar 监视 Prometheus 的 TSDB 目录,并把每个两小时的块上传到 RustFS 的 `thanos-data` 存储桶。Store Gateway 读取同一存储桶并回答针对历史块的查询,Query 因此既能通过 sidecar 解析实时数据,也能通过 Store Gateway 解析旧数据。 + +## 1. 创建项目文件 + +创建工作目录: + +```bash +mkdir rustfs-thanos +cd rustfs-thanos +``` + +创建环境文件,并替换两个凭证占位符: + +```ini title=".env" +RUSTFS_ACCESS_KEY= +RUSTFS_SECRET_KEY= +``` + +请为 `thanos-data` 存储桶使用专用凭证。不要将 `.env` 提交到版本控制。 + +创建 Prometheus 配置,并设置外部标签——Thanos 依赖它对块去重: + +```yaml title="prometheus.yml" +global: + scrape_interval: 5s + external_labels: + monitor: rustfs-demo + +scrape_configs: + - job_name: prometheus + static_configs: + - targets: ["localhost:9090"] + - job_name: rustfs + metrics_path: /metrics + static_configs: + - targets: ["rustfs:9000"] +``` + +创建 Thanos 对象存储配置: + +```yaml title="bucket.yml" +type: S3 +config: + bucket: thanos-data + endpoint: rustfs:9000 + access_key: ${RUSTFS_ACCESS_KEY} + secret_key: ${RUSTFS_SECRET_KEY} + insecure: true +``` + +Thanos 自身不会解析 `.env` 文件。启动前,把占位符替换为与 `.env` 相同的值: + +```bash +sed -i.bak "s|\${RUSTFS_ACCESS_KEY}|$(grep RUSTFS_ACCESS_KEY .env | cut -d= -f2)|;s|\${RUSTFS_SECRET_KEY}|$(grep RUSTFS_SECRET_KEY .env | cut -d= -f2)|" bucket.yml +``` + +创建 Compose 文件: + +```yaml title="compose.yaml" +services: + rustfs: + image: rustfs/rustfs-x86-musl:v2.3.1 + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + RUSTFS_VOLUMES: /data + RUSTFS_ADDRESS: ":9000" + RUSTFS_CONSOLE_ADDRESS: ":9001" + RUSTFS_CONSOLE_ENABLE: "true" + volumes: + - rustfs-data:/data + ports: + - "9000:9000" + - "9001:9001" + healthcheck: + test: ["CMD", "curl", "-sf", "http://127.0.0.1:9000/health"] + interval: 10s + timeout: 5s + retries: 6 + start_period: 10s + networks: + - thanos + + create-bucket: + image: rustfs/rc:latest + depends_on: + rustfs: + condition: service_healthy + environment: + RUSTFS_ACCESS_KEY: ${RUSTFS_ACCESS_KEY} + RUSTFS_SECRET_KEY: ${RUSTFS_SECRET_KEY} + entrypoint: + - /bin/sh + - -c + - | + /usr/bin/rc alias set rustfs http://rustfs:9000 "$${RUSTFS_ACCESS_KEY}" "$${RUSTFS_SECRET_KEY}" + /usr/bin/rc mb --ignore-existing rustfs/thanos-data + networks: + - thanos + + prometheus: + image: prom/prometheus:v2.53.1 + command: + - --config.file=/etc/prometheus/prometheus.yml + - --storage.tsdb.path=/prometheus + - --storage.tsdb.min-block-duration=2h + - --storage.tsdb.max-block-duration=2h + - --web.enable-lifecycle + volumes: + - ./prometheus.yml:/etc/prometheus/prometheus.yml:ro + - prom-data:/prometheus + ports: + - "9090:9090" + networks: + - thanos + + sidecar: + image: thanosio/thanos:v0.37.2 + command: + - sidecar + - --tsdb.path=/prometheus + - --prometheus.url=http://prometheus:9090 + - --objstore.config-file=/etc/thanos/bucket.yml + volumes: + - ./bucket.yml:/etc/thanos/bucket.yml:ro + - prom-data:/prometheus + depends_on: + create-bucket: + condition: service_completed_successfully + networks: + - thanos + + store: + image: thanosio/thanos:v0.37.2 + command: + - store + - --objstore.config-file=/etc/thanos/bucket.yml + - --data-dir=/data + volumes: + - ./bucket.yml:/etc/thanos/bucket.yml:ro + - store-data:/data + depends_on: + create-bucket: + condition: service_completed_successfully + networks: + - thanos + + query: + image: thanosio/thanos:v0.37.2 + command: + - query + - --http-address=0.0.0.0:9090 + - --store=sidecar:10901 + - --store=store:10901 + ports: + - "9091:9090" + depends_on: + - sidecar + - store + networks: + - thanos + +networks: + thanos: + +volumes: + rustfs-data: + prom-data: + store-data: +``` + +`--storage.tsdb.min-block-duration` 和 `--storage.tsdb.max-block-duration` 两个参数用于禁用 Prometheus 压缩。如果 TSDB 会本地压缩块,sidecar 会拒绝上传,因为本地块将不再与已上传的块一致。 + +## 2. 校验并启动部署 + +启动容器前先解析 Compose 文件: + +```bash +docker compose config +``` + +启动整个栈,等待 sidecar 报告就绪: + +```bash +docker compose up -d +docker compose logs sidecar | grep -m1 "status=ready" +``` + +确认 sidecar 已读取 Prometheus 的外部标签: + +```bash +docker compose logs sidecar | grep "external labels" +``` + +Thanos Query UI 在 `http://localhost:9091` 上提供服务,RustFS 控制台位于 `http://localhost:9001`。 + +## 3. 上传一个块到 RustFS + +sidecar 会在 Prometheus 压缩出块时上传,压缩发生在两小时的块边界上。要立即生成一个块,可通过管理 API 对 TSDB 做快照——在禁用压缩的情况下,sidecar 会直接上传 head 块的快照: + +```bash +curl -s -XPOST http://localhost:9090/api/v1/admin/tsdb/snapshot | head -c 200 +``` + +等待上传完成,然后查看 Prometheus 内的 shipper 状态: + +```bash +sleep 60 +docker compose exec prometheus cat /prometheus/thanos.shipper.json +``` + +`uploaded` 列表中应出现一个块 ID: + +```json +{ + "version": 1, + "uploaded": [ + "01M31EPTZC5E0SETZTP0SPFY79" + ] +} +``` + +## 4. 从 RustFS 查询历史数据 + +Store Gateway 会周期性同步存储桶。确认它已下载刚上传的块: + +```bash +docker compose logs store | grep "loaded new block" +``` + +通过 Query 前端在该块的时间范围内查询一条序列: + +```bash +START=$(date -u -d '2 hours ago' +%s) +END=$(date -u +%s) +curl -s "http://localhost:9091/api/v1/query_range?query=up%7Bjob%3D%22prometheus%22%7D&start=$START&end=$END&step=30" | head -c 300 +``` + +Store Gateway 用它从 RustFS 下载的块来响应请求,而 sidecar 负责实时的 head 数据——两条路径都经由同一个 Query 端点解析。 + +## 5. 在 RustFS 中验证对象 + +列出存储桶: + +```bash +docker compose exec rustfs /usr/bin/rc ls local/thanos-data/ -r +``` + +每个块保存为三个对象——chunk 文件、index 和 `meta.json`: + +```text +01M31EPTZC5E0SETZTP0SPFY79/chunks/000001 +01M31EPTZC5E0SETZTP0SPFY79/index +01M31EPTZC5E0SETZTP0SPFY79/meta.json +``` + +![RustFS 控制台中存储的 Thanos 块](./images/rustfs-thanos-blocks.png) + +## 6. 停止或重置部署 + +停止容器并保留所有数据: + +```bash +docker compose down +``` + +RustFS 卷会保留已上传的块,重启后 Store Gateway 依然可以回答历史查询。若要删除包括 RustFS 中块在内的所有数据,请追加 `--volumes`。 + +## 故障排查 + +### sidecar 日志出现 `Compaction needs to be disabled` + +Prometheus 必须以 `--storage.tsdb.min-block-duration` 等于 `--storage.tsdb.max-block-duration` 的方式运行——如 Compose 文件所示,两者都设为 `2h`。否则 sidecar 无法保证本地块保持不变,会拒绝上传。 + +### `The specified bucket does not exist` + +Thanos 不会创建存储桶。检查 `create-bucket` 服务是否成功完成: + +```bash +docker compose logs create-bucket +``` + +### 查询不到历史数据 + +确认 Store Gateway 至少加载了一个块(`docker compose logs store | grep "loaded new block"`),并且查询时间范围落在已上传块的窗口内——可在 RustFS 控制台中查看块的 `meta.json` 里的 `minTime` 和 `maxTime`。 + +## 后续步骤 + +- 在采用更多 Thanos 组件之前,请查阅 [S3 兼容性说明](/administration/protocols/s3)。 +- 通过[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Thanos 文档](https://thanos.io/tip/thanos/getting-started.md)添加 Compactor、Ruler 或 Receive,搭建生产拓扑。