The observation stream
metrics.observer is called synchronously from the check path with one CheckObservation:
onError to see the failure); a slow observer is your own latency, so never do I/O in one.
Pass an array to fan out to several sinks at once.
OpenTelemetry
otelMetricsObserver writes checks into a Meter. Alfiz does not depend on @opentelemetry/api. The adapter is typed structurally, so a real meter satisfies it with no cast:
Principals are omitted by default: unbounded cardinality, and PII-adjacent. Turn them on with
principals: true only if you know your backend can take it.
The same shape works for StatsD, Prometheus, or a log line. The observer interface is the extension point; the OpenTelemetry adapter is just the one that ships.
Sampling for high-traffic paths
sampleRate is evaluated with one Math.random() inside the call, before any observation is built. An unsampled check costs a comparison and allocates nothing, with no storage read and no coordination.
can or require maps one-to-one onto a user action and is usually worth keeping whole. A single server render fires hundreds of canAny, holds, and heldKeys checks, and one in fifty says the same thing.
A plain number applies to every shape. Every observation carries the rate that kept it, so counts extrapolate honestly: the aggregator and the OpenTelemetry adapter both report observed (what was seen) next to estimated (what it stands for), and never silently substitute one for the other.
Sampling decides only whether a check is counted. It can never change an answer, and the sampled path and the unsampled path evaluate identically.
Reading metrics directly
createMetricsAggregator() folds the stream into fixed memory and hands you the current window whenever you ask. That is a complete metrics API with no external system and no storage:
- Scope instances aggregate to scope type.
docs.doc:123counts asdocs.doc. Opt specific types in withscopeInstances: ["docs.folder"]when you want per-folder numbers; do not opt in a per-document type. - Principals are bounded. Exact distinct counts up to a cap, then an
overflowedflag, and never an unbounded set. - Every counter map is capped, and each batch reports how many observations a cap refused with
dropped. - Windows are tagged with an instance id and bounds, so batches from many app servers merge.
Storing usage, and the revocation safeguard
To answer “what breaks if I revoke this”, the counts have to outlive the process. Turn on storage in the Application (one extra table) and wire the client to it:Why it counts sole matchers
The obvious metric overwarns. A check satisfied by two grants loses nothing when one of them is revoked, so “this grant matched 1,200 checks” against a row fully shadowed by a broader grant is a warning that means nothing, and administrators learn to click through it. Alfiz counts both, and warns on the second:matched. The grant participated in an allow.soleMatch. The grant was the only row allowing. Revoking it would have flipped exactly this many checks to deny.
What the warning never says
When a grant shows no recorded use, Alfiz says exactly that and stops:No recorded use in the last 7 days. Absence of recent use is not evidence that removing this is safe: usage lags, metrics are sampled and lossy, and break-glass access is precisely the kind that goes unused for long stretches.Metrics tell you when a grant is load-bearing. They never tell you a grant is safe to remove, and copy you write on top of them should not either.
Usage numbers are counts, not audit. They are sampled, lossy, and dropped under back-pressure. Never derive an authorization decision or a compliance record from them. That is what
explain() and the audit log are for.Nothing joins your request path
Every part of this is built so that a metrics failure is a lost count and never anything more:- Observers are called synchronously but fire-and-forget; a throw is caught and the decision returns unchanged.
- Attribution is computed only when a check is sampled, from data the evaluation already produced.
- Delivery to storage is batched, pre-aggregated, and never awaited by a check, a render, or a write.
- Under back-pressure the sink drops batches instead of growing a queue. The right failure mode for a counter is losing counts; adding latency is not.
Where the data lives
Metrics stay in the deployment that produced them. The client hands batches to your own Application, and nothing carries them further: there is no uplink and no central metrics store. If you use the hosted dashboard, it can read your numbers for your own administrators, the same relayed read as every other admin surface, but Alfiz Cloud keeps no copy, and check volume is not a billing dimension in any tier. If you want the numbers somewhere central, the observer stream is how you send them there, to the destination you choose.Next steps
Metrics API
Every option, type, and method: observations, sampling, the aggregator, the OpenTelemetry adapter, and the usage reads.
Grants and revokes
The rows metrics attribute to, and the precedence rules behind
soleMatch.