2026-10-09 15:23:16 +02:00
2026-10-09 21:37:52 +02:00
2026-10-10 15:22:21 +02:00
2026-10-09 15:23:16 +02:00
2026-10-09 21:37:52 +02:00
2026-10-09 15:23:16 +02:00

LunarSky ObjectStore banner

ObjectStore

Small S3-compatible object storage for LunarSky.

Source: GitHub · Gitea mirror

Development software. Do not use it for production data. It is provided as is, without warranty, under the MIT License.

Stage: development License: MIT Runtime: Java 21 S3 API: partial support CodeFactor

Contents

Capability checklist

ObjectStore creates the configured default bucket at startup. Additional buckets share the configured capacity limit.

  • ✅ Persistent single-node storage with checksum verification
  • ✅ Optional dm-crypt-backed data mounts with startup guards
  • ✅ Configurable per-object and total logical size limits
  • ✅ CLI status, version, and full payload verification
  • ✅ Local cluster prototype with stable node IDs and host-aware placement code
  • ✅ Optional HTTPS between cluster processes and storage nodes
  • ✅ Opt-in automatic repair, rebalance, and guarded garbage collection in the local cluster
  • ✅ Metadata backup and tested restore to a separate local PostgreSQL instance
  • ✅ Manual two-machine durability and metadata-restore drill
  • ⬜ Production multi-server deployment and metadata failover

New objects retain their content type and key. Objects written by the earlier single-node format remain readable, but cannot appear in listings until overwritten because their original keys were not stored.

S3 API support checklist

  • ✅ Header-based and presigned-query AWS Signature Version 4 authentication
  • ✅ Path-style and configurable virtual-hosted-style bucket addresses
  • ✅ PutObject, GetObject, HeadObject, and DeleteObject in both modes
  • ✅ Single-range GET and ListObjectsV2 in both modes
  • ✅ SHA-256 payload verification and x-amz-checksum-sha256 in both modes
  • ✅ Content-MD5 and CRC32, CRC32C, CRC64NVME, XXHash64, XXHash3, XXHash128, SHA-1, SHA-256, SHA-512, and MD5 checksum headers on PutObject and UploadPart
  • ✅ CreateMultipartUpload, UploadPart, CompleteMultipartUpload, and AbortMultipartUpload in both modes
  • ✅ ListParts and ListMultipartUploads in both modes
  • ✅ Presigned URLs and signed streaming Signature V4 uploads, including signed checksum trailers
  • ✅ CreateBucket, HeadBucket, DeleteBucket, and ListBuckets in both modes
  • ✅ CopyObject across owned buckets, with COPY and REPLACE content-type and user-metadata behavior
  • ✅ User metadata on object and multipart uploads; object tags on upload, copy, and the tagging subresource
  • ✅ CRC64NVME and XXHash checksums, checksum trailers, and persisted object checksum metadata
  • ✅ Bucket versioning with retained versions, delete markers, and version-specific reads, copies, deletes, and tags in both modes
  • ✅ Basic bucket and object ACL grants, public reads, and multiple access-key identities
  • ⬜ Full AWS ACL ownership controls, email grantees, and bucket policies

This is an S3 API subset, not full AWS S3 compatibility. Unsupported S3 operations and Amazon-specific headers are rejected. Buckets use the same three-to-63-character lowercase names as the configured default bucket. The default bucket cannot be deleted through the API. Bucket creation currently accepts the empty-body request used for the configured region. Tags do not control access. Versioning supports enabled and suspended states, historical object versions, null versions, and delete markers. Retained versions count toward the capacity limit. Deleting a specific version is permanent. Lifecycle expiration, MFA Delete, and ListObjectVersions delimiter grouping are not supported yet. User metadata values currently accept printable ASCII only; non-ASCII metadata header encoding is not yet supported. Checksums supplied with PutObject are validated before publication and retained across restarts, copies, and object versions. When no checksum is supplied, ObjectStore calculates and stores CRC64NVME, including for completed multipart uploads. Request x-amz-checksum-mode: ENABLED on GetObject or HeadObject to receive the stored checksum. UploadPart checksums are validated but are not yet returned by ListParts or combined into a multipart checksum; the completed object's CRC64NVME covers its full content. Presigned URLs use query Signature V4 with a maximum seven-day expiry. Streaming uploads support signed aws-chunked payloads and one signed checksum trailer. Temporary credentials and other streaming payload modes remain unsupported. Copies use the existing object size limit.

Single-node setup

Requires Docker Compose. Copy .env.example to .env, then set S3_ACCESS_KEY and S3_SECRET_KEY to unique values. The access key must be at least 16 alphanumeric characters; the secret must be at least 32 characters.

cp .env.example .env
docker compose up -d --build
curl http://127.0.0.1:9000/health

Port 9000 binds to localhost. Data stays in the object-data Docker volume. docker compose down -v deletes that volume.

The standalone defaults are 128 MiB per object and 2 GiB total. Set MAX_OBJECT_BYTES and MAX_TOTAL_BYTES in .env to change them. Incomplete multipart uploads consume space until aborted.

Encrypted storage

At-rest encryption is an opt-in deployment setting backed by a host-mounted dm-crypt/LUKS filesystem. ObjectStore does not encrypt bytes itself or store an encryption key in its environment. Prepare and unlock the LUKS filesystem outside Docker, then set ENCRYPTED_STORAGE_ROOT to its exact mount point and ENCRYPTED_VOLUME_ID to a stable 32-character lowercase hex identifier in your private environment file. Generate the identifier with openssl rand -hex 16; it is a mount guard, not a cryptographic key. Keep the LUKS recovery key separately from the data and backups.

With the encrypted filesystem mounted, prepare empty directories using sh scripts/prepare-encrypted-storage.sh single /path/to/objectstore.env or sh scripts/prepare-encrypted-storage.sh cluster /path/to/cluster.env. The script reads only the two encryption settings from that file; it does not execute it. It requires an active dm-crypt-backed mount at ENCRYPTED_STORAGE_ROOT and refuses to mark a nonempty directory. Run it with permission to create the directories and set container ownership. Enable the matching Compose override:

docker compose --env-file /path/to/objectstore.env -f compose.yaml -f compose.encrypted.yaml up -d --build
docker compose --env-file /path/to/cluster.env -f compose.cluster.yaml -f compose.cluster.encrypted.yaml up -d --build

Use the first command for single-node mode or the second for the local development cluster, not both. The overrides bind encrypted directories for object data, cluster nodes, PostgreSQL, and cluster staging files. Each process checks a marker stored on its mounted directory before opening data; PostgreSQL checks before starting. If a mount is missing after reboot, the service refuses to start instead of silently creating a new plaintext store. Provision an unlock-and-mount service before starting Compose; a container restart cannot unlock LUKS. The preparation script checks the host mount, while the runtime marker is only a guard against an absent or wrong mount.

Existing Docker volumes are not migrated automatically. Stop the stack, take and verify a backup, prepare the empty encrypted directories, then copy the stopped volumes into their corresponding directories before starting with the override. For PostgreSQL, copy the existing data directory into cluster/metadata/pgdata because the encrypted override changes PGDATA. Verify the restored objects and database before retiring the original volumes. The cluster overlay places all local test containers on one encrypted host; a future multi-host deployment needs an encrypted data mount and key recovery plan on each host.

This covers the specified data and staging mounts, but not Docker logs, host swap, other temporary files, external backups, or a compromised running host. Encrypt those separately as appropriate, use TLS for network paths, restrict access, and test backup recovery. Encryption at rest is one security measure; enabling it does not by itself establish GDPR compliance.

Access keys and ACLs

S3_ACCESS_KEY is the owner identity. Its secret is S3_SECRET_KEY. Additional keys are optional: place one ACCESSKEY:secret pair per line in a file mounted read-only inside the container, and set S3_CREDENTIALS_FILE=/run/secrets/s3-credentials in the environment file. Use a Compose override to mount an absolute host path at /run/secrets/s3-credentials for the objectstore service (or gateway in cluster mode). Each access key needs 16–128 alphanumeric characters and each secret at least 32 characters. A restart loads changes to that file. Keep it outside Git and protect it as a secret. Additional identities have no access until the owner grants it.

The owner can send x-amz-acl: public-read or signed x-amz-grant-* headers on bucket creation, object uploads, copies, and multipart initiation. GET and PUT ?acl support bucket and object ACLs; a PUT accepts either signed grant headers with an empty body or an XML AccessControlPolicy. A grantee ID is its configured access key. Supported permissions are READ, WRITE, READ_ACP, WRITE_ACP, and FULL_CONTROL. The AllUsers group is limited to READ; AuthenticatedUsers is also recognized. Anonymous reads work only where AllUsers has a read grant. Replacing an object starts with a private ACL unless the new upload supplies grants; older version ACLs remain attached to their versions.

This is a deliberately limited ACL subset. The configured owner owns every bucket and object. A WRITE_ACP grantee can change an object ACL only on a retained, non-null version; updates to current unversioned and null versions require the owner key. List-versions, multipart inspection, tagging, versioning controls, bucket creation/deletion, and the capability endpoint remain owner-only. There are no IAM policies, email grantees, Object Ownership modes, Block Public Access settings, or temporary credentials. Granting public bucket READ exposes object names through ListObjectsV2; granting public object READ exposes that object's bytes. Review those grants before exposing a gateway to the internet.

CLI and tests

From the server shell, run the CLI inside the running container from the directory containing compose.yaml:

docker compose exec objectstore objectstore status
docker compose exec objectstore objectstore verify
docker compose exec objectstore objectstore version

The startup log shows the LunarSky banner, version, and a small storage summary (docker compose logs --tail=20 objectstore). status reports object and multipart usage. verify also checks stored payload hashes and exits nonzero on an error. Both commands can run while the service is live; they are not a snapshot or a backup.

Capability discovery

ObjectStore provides a signed GET /_objectstore/capabilities endpoint. Use the same header-based Signature V4 authentication as the S3 API. It returns JSON with schemaVersion: 1, the service and software version, storage mode, supported operation names, and configured limits. The endpoint does not disclose credentials or cluster topology. The response has Cache-Control: no-store because limits can change after a restart.

operations lists implemented API operations, not every AWS option for each operation or the caller's authorization to use them. limits.maxObjectBytes applies to a completed object and to each uploaded part; limits.maxTotalBytes is the logical storage limit; limits.maxParts is 10,000. Future schema version 1 responses may add fields. Clients should ignore unknown fields and treat unknown operation names as unsupported by their own implementation. This endpoint is an ObjectStore extension; a missing endpoint on another S3-compatible service does not prove that a feature is unavailable.

Run sh scripts/test.sh with JDK 21 to test from source. See TESTS.md for coverage, the disposable Docker cluster suite, and the limits of those tests.

The server includes hash4j for streaming XXHash checksums. Its Apache-2.0 license is included at lib/LICENSE.hash4j.

Java client

The JDK-only Java client works with ObjectStore and other S3-compatible endpoints. It supports object transfers, listing, multipart uploads, and read-only capability discovery. An AWT image manager provides a small desktop example. Run bash client/scripts/build.sh to produce its JAR and Javadoc locally.

Local cluster prototype

The local cluster prototype starts three segment containers and one PostgreSQL container on the same Docker host. Copy .env.cluster.example to a private environment file, replace all five credential values, and run:

docker compose --env-file /path/to/cluster.env -f compose.cluster.yaml up -d --build
sh scripts/test-cluster.sh /path/to/cluster.env
docker compose --env-file /path/to/cluster.env -f compose.cluster.yaml run --rm repair

The cluster S3 endpoint binds to 127.0.0.1:9001; storage nodes and PostgreSQL have no published ports. The separate repair container holds the repair credential and restores missing or corrupt replicas.

Multipart parts are stored on cluster nodes and indexed in PostgreSQL. Incomplete uploads count toward the logical capacity limit; abort them to release that capacity. Repair includes staged parts. The gateway upgrades the metadata schema when it starts, so back up the database before upgrading an existing cluster.

Node UUIDs persist on their volumes, and replica manifests use those UUIDs so reordering configured URLs cannot move an existing replica. Each node also has an operator-assigned physical host UUID. New writes require acknowledgements from two different host UUIDs. The optional CLUSTER_TEST_NODE_DOMAINS=true override counts containers instead, solely for local process tests; all containers in this Compose file share one physical host.

Node transport TLS

The default local Compose cluster uses HTTP inside its private Docker network. For an HTTPS test, give each node a PKCS#12 keystore containing its private key and a certificate whose DNS subject alternative name matches its CLUSTER_NODES hostname. Give the gateway, repair, garbage collection, and maintenance processes a PKCS#12 truststore containing the issuing CA or each node certificate. Mount the files read-only and keep the keystores and password files outside Git.

Set NODE_TLS_KEYSTORE and NODE_TLS_PASSWORD_FILE on each node. Set CLUSTER_TLS_TRUSTSTORE and CLUSTER_TLS_PASSWORD_FILE on every process that contacts nodes, and change each node URL to https://. With a truststore configured, HTTP node URLs are rejected. The client verifies the certificate chain and hostname; a failed handshake does not fall back to HTTP. Both settings in each pair are required. Restart affected processes after rotating certificates or truststores.

The optional Compose TLS overlay expects node-a.p12 through node-d.p12, matching .pass files, and trust.p12 with trust.pass in CLUSTER_TLS_DIR. Set that variable to a private certificate directory and include both Compose files:

CLUSTER_TLS_DIR=/private/objectstore-certs docker compose --env-file /path/to/cluster.env \
  -f compose.cluster.yaml -f compose.cluster.tls.yaml up -d --build

The files must be readable by container UID 10001 without making private keys or passwords world-readable. Add HTTPS URLs for additional nodes when expanding the cluster.

This secures node traffic only. The local cluster still lacks automatic PostgreSQL failover, database TLS configuration, encryption at rest, and production multi-server validation. Its HTTP S3 gateway remains bound to localhost; use a separate trusted proxy for external TLS. Do not treat the TLS overlay as a production deployment.

Metadata primary routing

POSTGRES_JDBC_URL can override the database URL for the gateway, repair, garbage collection, and maintenance processes. Its local Compose default now uses targetServerType=primary, and /ready returns unavailable when the database is read-only or in recovery. The base Compose file still starts and waits for its own single PostgreSQL container; it is not an HA deployment. In an independently managed deployment, list the PostgreSQL hosts in the JDBC URL and keep targetServerType=primary:

jdbc:postgresql://db-a:5432,db-b:5432/objectstore?targetServerType=primary&connectTimeout=3&socketTimeout=10

ObjectStore opens a new database connection for each operation, so the JDBC driver can select a promoted primary after the old one is stopped. This does not promote a standby, fence the old primary, configure synchronous replication, or guarantee that a recently acknowledged write reached the standby. Those jobs belong to a separately operated PostgreSQL HA system. Never allow two writable metadata databases: they can diverge while serving different ObjectStore requests. A request interrupted during failover has an uncertain outcome; verify it before retrying a non-idempotent operation. For remote database connections, configure PostgreSQL TLS and use JDBC sslmode=verify-full with a mounted CA certificate. The provided local Compose database does not enable TLS.

Migrating a local cluster

For an existing local three-node cluster that stores replicas by URL position:

  1. Stop the old gateway. Back up PostgreSQL and every node volume, and record the original ordered CLUSTER_NODES.
  2. Start the nodes with the new image and a separate CLUSTER_REPAIR_TOKEN. Do not start the new gateway yet.
  3. Run objectstore cluster-migrate --check in a one-off gateway container. Compare legacy_0 through legacy_2 and their URLs with the pre-upgrade inventory.
  4. Run objectstore cluster-migrate --apply id0,id1,id2 with those node IDs in the original order. The command verifies every listed live replica before conversion.
  5. Start the new gateway only after the command reports cluster_format=2. Do not restart an old gateway against the converted database.

The old format did not record the original URL mapping, so the inventory check is essential. The isolated migration fixture tests conversion; building this code does not automatically upgrade the running local prototype.

Adding a cluster node

objectstore cluster-join http://new-node:9100 expected-host-uuid registers an additional local node. Add its URL to CLUSTER_NODES and restart the gateway. A repair pass then copies existing segments to their preferred nodes and removes obsolete replicas from the manifest only after verifying the replacements. scripts/test-cluster.sh exercises a fourth container joining, receiving new segments, and rebalancing existing ones.

Cluster maintenance and recovery

From the Compose directory, use docker compose --env-file /path/to/cluster.env -f compose.cluster.yaml as the command prefix:

docker compose --env-file /path/to/cluster.env -f compose.cluster.yaml --profile automatic up -d maintenance
docker compose --env-file /path/to/cluster.env -f compose.cluster.yaml run --rm gc
sh scripts/backup-cluster-metadata.sh /path/to/cluster.env /path/to/metadata.dump

The maintenance service is opt-in. It repairs missing or corrupt replicas and rebalances them every 60 seconds by default. Set CLUSTER_MAINTENANCE_INTERVAL_SECONDS to change the interval. The gc command is a dry run; use gc --apply only after checking its candidate count and keeping independent backups. Cleanup records each orphan on one pass and waits at least CLUSTER_GC_MIN_AGE_SECONDS before deleting it on a later pass. The default age is 14 days. To permit deletion, set CLUSTER_BACKUP_RETENTION_SECONDS to your actual backup retention in seconds; it must be at least one day and shorter than the cleanup age. Scheduled cleanup also requires CLUSTER_GC_ENABLED=true on the maintenance service. The backup command writes a verified PostgreSQL archive with private file permissions. Restore it to a separate database and point a gateway at that database only after validating the restore. A backup restore is manual recovery, not automatic failover.

Limits and safety

Public-client limits are disabled by default. The localhost Compose setup is unchanged. For an endpoint that accepts external clients, set one or both of PUBLIC_REQUESTS_PER_SECOND and PUBLIC_BYTES_PER_SECOND to a positive number in .env. The first limits requests per client IP with a token bucket; the second paces upload and download bytes through one shared per-IP budget. PUBLIC_REQUEST_BURST and PUBLIC_BYTE_BURST default to one second of their respective rates. PUBLIC_MAX_IN_FLIGHT_PER_IP defaults to 8 when either rate is enabled. Requests above the rate or concurrency limit receive S3 503 SlowDown and Retry-After: 1; an admitted transfer is paced rather than cut off. These are per-gateway limits, not cluster-wide quotas.

MAX_IN_FLIGHT_REQUESTS controls the per-gateway concurrency cap (default 16, range 1–1024); an additional request receives 503 SlowDown. HTTP_BACKLOG controls the listening socket backlog (default 64, range 1–4096). Raise either only after a mixed upload/download load test, because each active request can hold memory, disk space, and database connections. The configured object and total storage limits still apply. MAX_OBJECT_BYTES currently cannot exceed 1 GiB.

Set S3_VIRTUAL_HOST_SUFFIX to a DNS suffix such as s3.example.com to accept bucket.s3.example.com/key alongside path-style /bucket/key. Configure DNS and TLS for the bucket hostnames, and preserve the client's original Host header through the proxy; Signature V4 signs it. The suffix is empty by default. This changes request parsing, not bucket naming rules or DNS configuration.

For example, to start with 100 requests per second and 16 MiB/s combined upload and download per IP, set PUBLIC_REQUESTS_PER_SECOND=100 and PUBLIC_BYTES_PER_SECOND=16777216. Both start with a one-second burst. Adjust these numbers after measuring the actual workload; do not copy them as a universal production policy.

ObjectStore uses the socket peer as the client IP and ignores forwarded-IP headers by default. If a reverse proxy sits between external clients and ObjectStore, set PUBLIC_TRUSTED_PROXY_IPS to the exact IP address that ObjectStore sees for that proxy and configure the proxy to replace X-Real-IP with its actual client IP. A request from that trusted peer without exactly one valid numeric X-Real-IP is rejected. Do not trust an address reachable by arbitrary clients, and preserve the original Host header for Signature V4. In Docker, the proxy's address as seen by the container may be a bridge address rather than 127.0.0.1. If proxy trust is not configured, all clients behind that proxy share its budget; this is safe from header spoofing but may throttle them together.

Enable these limits only on a deliberately public endpoint. They also apply to direct localhost storage calls to that same endpoint; leave them disabled for the local-only setup or run a separate local-only instance if local storage calls must be exempt. A direct loopback /health probe without a forwarded-IP header remains exempt. The byte limit is aggregate ingress plus egress for each IP, and several users behind one NAT share it. If a proxy buffers complete uploads before forwarding them, the upload byte limit controls proxy-to-ObjectStore traffic, not the client's initial upload speed; disable request buffering when end-to-end upload pacing is required. It is a fairness control, not a defense against connection floods before the Java handler runs. Put an internet-facing proxy or firewall in front of the gateway for TLS, connection limits, request timeouts, and buffering controls. Do not expose storage nodes or PostgreSQL publicly.

The default local cluster still uses plaintext node and PostgreSQL connections. The optional TLS overlay secures node traffic, but PostgreSQL TLS, automatic metadata failover, scoped credentials, and physical host verification remain absent. Garbage collection can remove data required by an older metadata backup, so its retention guard is essential. Host UUIDs are operator labels, not proof that machines have separate power, disks, or network paths. Keep CLUSTER_LOCAL_DEV=true limited to local tests.

The standalone cluster node binds to localhost by default. Set NODE_BIND only for a private test network; the Compose file binds inside its private Docker network. PostgreSQL JDBC 42.7.14 is bundled in the image with its license inside the JAR.

Disclaimer

ObjectStore is independent software. It is not affiliated with, sponsored by, or endorsed by Amazon or Amazon Web Services (AWS). “S3-compatible” describes only the API subset listed above.

Keep independent backups; data loss is possible. The software is provided “as is” under the MIT License. To the extent permitted by applicable law, the authors and maintainers are not responsible for data loss or other damages arising from its use.

AI contributions

AI tools assisted with parts of this project. Maintainers review releases and remain responsible for what ships.

S
Description
Small S3-compatible object storage for LunarSky
Readme MIT
1.5 MiB
1 Stars 1 Watchers 0 Forks
Languages
Java 93.6%
Python 3.7%
Shell 2.4%
Dockerfile 0.3%