Simple to start · platform-grade to grow
Turn Kubernetes state into durable inventory. Declare what matters once. Kollect keeps it current and delivers it to Git, object storage, databases, and event streams. Select resources by GVK, extract the attributes you need with CEL or JSONPath, and every sink receives the same canonical rows in parallel.
Start with one sink. Grow to a whole platform. A single pipeline can write an inspectable Git
history, a queryable database record, an object-store snapshot, or an event stream—without scripts
or API-server hammering. As adoption grows, nothing gets rebuilt: the same rows fan out to more
sinks, and KollectScope keeps it multi-tenant. Every team owns its inventory as configuration,
not code, in its own namespace; consumers read export data, never unbounded list/watch against
the live cluster.
Read the docs: platformrelay.github.io/Kollect — architecture, quick start, CR reference, ADRs, and examples. This README is the front door; the site is the map.
Pre-1.0. Kollect uses a
v1alpha1API. Breaking API or default changes may ship in minor releases before 1.0; release notes and migration guidance call them out. See the roadmap for current maturity.
- Decoupled read model — consumers query a sink, not the apiserver. No RBAC blast radius, no watch-storm risk, no etcd size limits (why).
- Event-driven, no polling — one shared informer per GVK keeps inventory current as the cluster changes (ADR-0301).
- Schema-flexible — declare the attributes you want in a
KollectProfile; no bespoke collector per resource kind. - Pluggable sinks, no privileged backend — the same snapshot fans out to Git, Postgres, object store, or an event stream (sink taxonomy).
- Multi-tenant by design —
KollectScopegates which teams, namespaces, and sinks each tenant may use. - Fleet-ready — N single-mode operators → one shared sink, partitioned by
spec.cluster; no central hub tier to operate (ADR-0501). - Scale-aware architecture — shared informers, export sharding, and tunable reconcile/dispatch concurrency; the performance guide separates measured evidence from targets (performance).
A real pipeline is a handful of Kubernetes resources. The first-inventory walkthrough collects container images from Deployments and exports them to Git for an inspectable audit trail:
flowchart LR
Profile["<b>KollectProfile</b><br/>Deployment schema"]
Target["<b>KollectTarget</b><br/>select Deployments"]
Inv["<b>KollectInventory</b><br/>aggregate · debounce · export"]
Snap["<b>KollectSnapshotSink</b>"]
Db["<b>KollectDatabaseSink</b>"]
Ev["<b>KollectEventSink</b>"]
K8s[("Kubernetes API")]
Profile --> Target
K8s -- "informer per GVK" --> Target
Target --> Inv
Inv --> Snap
Inv --> Db
Inv --> Ev
Snap --> SnapOut["Git · GitLab · S3 · GCS"]
Db --> DbOut["Postgres · MongoDB"]
Ev --> EvOut["Kafka"]
Spin up the full pipeline on a local kind cluster in one command (needs Docker, kind, kubectl, and Task):
git clone https://github.com/platformrelay/kollect.git && cd kollect
task dev-up # build, create kind cluster, install operator + sample CRs
kubectl get kinv,ktgt,ksnap,kdb -A # watch the pipeline come uptask dev-up builds the manager, boots a kollect-dev kind cluster, installs the operator, and
applies the sample Profile → Sink → Target → Inventory pipeline. Watch the KollectInventory
Ready condition, then read your sink — the live demo repo
shows what the Git export looks like.
Full walkthrough — local evaluation and Helm install: Install Kollect →
flowchart LR
API["Kubernetes API"] -->|shared informers| Snapshot["Canonical inventory snapshot"]
Snapshot -->|debounce| Inventory["KollectInventory"]
Inventory --> SnapshotSinks["Git · GitLab · S3 · GCS"]
Inventory --> DatabaseSinks["Postgres · MongoDB · BigQuery"]
Inventory --> EventSinks["Kafka · NATS"]
The in-memory snapshot per inventory is canonical; every sink is a projection of it — no single backend is privileged (sink roles). Sinks are split into three CRD families (ADR-0414):
| Sink family | Examples | Good for |
|---|---|---|
KollectSnapshotSink |
Git, GitLab, S3, GCS | Audit, diff, GitOps-friendly history |
KollectDatabaseSink |
Postgres, MongoDB | Rich queries for portals and dashboards |
KollectEventSink |
Kafka, NATS | Change streams, downstream consumers |
Honest maturity tiers — see the roadmap for release timing.
| Family CRD | spec.type |
Status |
|---|---|---|
KollectSnapshotSink |
git |
Core — production-ready |
KollectSnapshotSink |
gitlab |
Core |
KollectSnapshotSink |
s3 |
Core |
KollectSnapshotSink |
gcs |
Beta — shipped, maturing |
KollectDatabaseSink |
postgres |
Core |
KollectDatabaseSink |
mongodb |
Beta |
KollectDatabaseSink |
bigquery |
Beta — analytics SQL |
KollectEventSink |
kafka |
Beta |
KollectEventSink |
nats |
Beta — JetStream emitter |
KollectSnapshotSink |
azureblob |
Planned — needs real backend (roadmap) |
KollectSnapshotSink |
s3, gcs with serialization.format: parquet |
Beta — shipped object-store output mode |
Full payload lives in sinks; CR .status holds summaries only (etcd limits).
Kollect is designed for large single clusters and multi-cluster fleets. The performance guide distinguishes reproducible results from design targets and documents tuning for reconcile concurrency, export debounce, and sharding. Fleet fan-in uses shared sinks rather than a hub merge tier.
| Topic | Link |
|---|---|
| Problem statement, CRD model, reconciliation | Architecture |
| Locked platform decisions | Platform decisions |
| CR fields, RBAC, failure modes | CR reference |
| Multi-cluster fleet | ADR-0501 |
| Sink taxonomy (state vs stream) | ADR-0401 |
| Shipped, next, and later work | Roadmap |
| Examples index | Examples |
| Example: Deployment → Git export | Walkthrough |
| Live demo inventory (Git sink) | kollect-inventory-demo |
Developers: run task lint, task test, and task verify before opening a PR —
CONTRIBUTING.md.
| Contributing | CONTRIBUTING.md — DCO, PR workflow, good first tasks |
| Code of Conduct | CODE_OF_CONDUCT.md — Contributor Covenant v2.1 |
| Governance | GOVERNANCE.md — roles, decisions, continuity |
Report vulnerabilities privately — see SECURITY.md. Security architecture: docs/ASSURANCE-CASE.md.
Copyright (c) 2026 Konrad Heimel. Licensed under the MIT License.