Deployment and release are two separate decisions.
8 domains · 24 rules.
DEPLOY · CONTROL · RELEASE
A feature flag evaluated at runtime against a remote configuration is fundamentally different from a hardcoded conditional. One lets you change system behavior without a deploy. The other is just a branch. Only one qualifies as infrastructure.
Flag evaluation on the hot path of a request must be a sub-millisecond local lookup against a cached config. A network call per flag evaluation adds latency, creates a hard dependency on the flag service, and will eventually cause an outage.
A controlled rollout is a sequence of stages, each validated with metrics before proceeding. The goal is to catch regressions when they affect 1% of users — not after they reach everyone.
user_id — not per request, not per session. A user re-evaluated randomly is simultaneously in both control and treatment.A kill switch is an operational flag that disables a feature or integration within seconds, without a deploy. It is not a feature flag with stricter semantics — it is a separate class of infrastructure with separate ownership, separate testing, and no expiry date.
Experiment flags share the same evaluation infrastructure as release flags but have fundamentally different lifecycle rules. An experiment without a pre-defined success metric and a hard end date is not an experiment — it is a permanent A/B split with no decision criteria.
The metric and the threshold that determine a winner must be agreed before the flag is turned on. Defining success after seeing results is HARKing — it produces the answer you wanted, not the answer that is true. Write the decision criteria in the flag description field.
"We'll run this for two weeks" is not experiment design. Required sample size is determined by the minimum effect size you care about and the baseline conversion rate. Running an underpowered experiment for two weeks produces inconclusive results regardless of how long you wait.
When an experiment concludes, the losing code path is removed — win or lose. Keeping dead variants in the codebase adds maintenance burden, makes the code harder to reason about, and turns temporary experiment flags into permanent conditions nobody understands.
Every active flag is a code path that must be maintained, tested (flag-on and flag-off), and reasoned about. Ten active release flags means your system has up to 1,024 possible behavioral states. Flag debt does not feel real until it is unmanageable.
A flag created without an agreed removal date will still be in the codebase two years later. At creation: set expected rollout completion date, name the owner of cleanup, and define what "done" looks like. This is non-negotiable for release and experiment flags.
Stale flags appear in sprint backlog and code review as explicitly as stale dependencies or TODO comments. The engineering team maintains a live inventory of all active flags: type, owner, age, and status. Flags older than their intended lifetime trigger a review.
A feature is done when the flag code path is removed, the config is deleted from the flag service, and the dead branch is cleaned up — not when it reaches 100% rollout. Cleanup is a tracked deliverable with the same weight as the feature itself.
Without flag-level observability, debugging inconsistent behavior is nearly impossible. "Why is this user not seeing the feature?" becomes a multi-hour investigation through targeting rules, SDK versions, and cache states. Evaluations must leave a trail.
user_id or tenant_id it was evaluated for. This is the only way to answer "why did user X get variant Y?" in under 30 seconds.Every flag introduces a code branch. A test suite that only exercises the default flag state will pass CI and break production when the flag is enabled. Testing with flags is not optional — it is part of the correctness contract of the feature.
Every test that exercises a flagged code path must run twice: once with the flag enabled, once disabled. Use flag overrides or test fixtures to set flag state explicitly. Never rely on default values in tests — defaults change.
Tests that resolve flags against a remote service are slow, flaky (network dependency), and couple test results to production flag configuration. Mock or override flag state locally in all test and CI environments. Flag evaluation in tests is a local operation.
A targeting rule that accidentally enables a feature for all tenants instead of beta tenants is a production incident. Targeting rule changes have the same blast radius as code changes affecting user-facing behavior — they require the same review and the same rollout discipline.
Deploy when engineering is ready.
Release when the business is ready.
Kill switches restore service without a rollback.
Flags accumulate debt. Schedule cleanup at creation.
A flag at 100% with no removal date is permanent technical debt.