Architecture Reference // 2026

Feature
Flagging

Deployment and release are two separate decisions.
8 domains  ·  24 rules.

CODE
→
FLAG
→
USER

DEPLOY  ·  CONTROL  ·  RELEASE

01 // Foundation

Flags are not if statements

A feature flag evaluated at runtime against a remote configuration is fundamentally different from a hardcoded conditional. One lets you change system behavior without a deploy. The other is just a branch. Only one qualifies as infrastructure.

Flags decouple deployment from release
Code can be deployed to 100% of servers and released to 0% of users. This decoupling removes the time pressure of deployment from the risk of exposure — you ship when engineering is ready, you release when the business is ready.
Four flag types — four distinct lifecycles
Release flags (new feature rollout, temporary), Experiment flags (A/B tests, time-bounded), Kill switches (operational control, permanent), Permission flags (tier/tenant access, permanent). Conflating them produces flags with no clear owner and no cleanup criteria.
The flag is not the feature — it is the gate
A flag at 100% rollout with no removal date is permanent technical debt. The feature is not done when it reaches 100% — it is done when the flag code path is removed, the config is deleted, and the dead branch is cleaned up.
Four flag types
Release — temporary
Experiment — time-bounded
Kill switch — permanent
Permission — permanent
02 // Architecture

Evaluation must be local, not remote

Flag evaluation on the hot path of a request must be a sub-millisecond local lookup against a cached config. A network call per flag evaluation adds latency, creates a hard dependency on the flag service, and will eventually cause an outage.

Evaluate server-side for anything authoritative
Client-side flag evaluation (SDK in the browser) is acceptable for UI gating. Any flag that controls business logic, rate limits, pricing, or access control must be evaluated server-side. Client-side flags are advisory — server-side flags are authoritative.
SDK caches config locally — no per-request network calls
Use a flag SDK that streams or polls config updates and evaluates locally. The flag service going down must not take your application down with it. Evaluation is a hash-map lookup; a network call is a dependency.
Default state must be explicit and safe
Every flag must define what happens when the flag service is unreachable or the flag config is missing. The default must be the safe behavior — feature off for kill switches, existing behavior for release flags. A flag system that throws on evaluation failure is worse than no flag system.
03 // Targeting & Rollout

0% to 100% is not a rollout strategy

A controlled rollout is a sequence of stages, each validated with metrics before proceeding. The goal is to catch regressions when they affect 1% of users — not after they reach everyone.

01
Start with internal users, not a percentage
The first rollout target is always employees and internal tooling. Internal users catch obvious issues — broken UI, missing data, wrong copy — before any customer is exposed. This stage costs nothing and has no blast radius.
02
Expand to beta / opt-in customers
A defined beta cohort — customers who have opted into early access — provides real-world signal under production conditions before a percentage rollout begins. Feedback from one willing beta customer is worth more than metrics from a random 1%.
03
Progress by percentage with metric gates
1% → 5% → 20% → 50% → 100%. At each stage, validate that error rate, latency, and conversion are within acceptable bounds before proceeding. Sticky assignment per user_id — not per request, not per session. A user re-evaluated randomly is simultaneously in both control and treatment.
04
Schedule flag removal before you reach 100%
When you flip a flag to 100%, the cleanup ticket should already exist and be assigned. GA is not done. GA plus flag removal is done. Teams that skip this step accumulate flag debt that compounds until it becomes unmanageable.
04 // Kill Switches

The fastest incident mitigation tool you have

A kill switch is an operational flag that disables a feature or integration within seconds, without a deploy. It is not a feature flag with stricter semantics — it is a separate class of infrastructure with separate ownership, separate testing, and no expiry date.

Every major feature and external integration needs one
Checkout, payments, email sending, external enrichment services, third-party SDKs — all need independently disableable kill switches. When a payment provider goes down, you need to disable that provider in seconds. Not a config deploy. Not a code rollback. One flag flip.
Kill switches must be tested outside of incidents
A switch discovered to be broken during a SEV1 is worse than no switch at all — it wastes time and creates false confidence. Test every kill switch in production on a scheduled basis during low-traffic windows. The test is: can a non-engineer flip this and verify the system behaves correctly?
Never include kill switches in the flag cleanup process
Release flags are temporary. Kill switches are permanent infrastructure. They are stored separately, documented separately, and explicitly excluded from flag debt reviews. Deleting a kill switch is a deliberate architectural decision, not routine maintenance.
05 // Experiments

A/B tests are not feature rollouts with extra steps

Experiment flags share the same evaluation infrastructure as release flags but have fundamentally different lifecycle rules. An experiment without a pre-defined success metric and a hard end date is not an experiment — it is a permanent A/B split with no decision criteria.

Define the success metric before enabling

The metric and the threshold that determine a winner must be agreed before the flag is turned on. Defining success after seeing results is HARKing — it produces the answer you wanted, not the answer that is true. Write the decision criteria in the flag description field.

Sample size is math, not calendar time

"We'll run this for two weeks" is not experiment design. Required sample size is determined by the minimum effect size you care about and the baseline conversion rate. Running an underpowered experiment for two weeks produces inconclusive results regardless of how long you wait.

Losing variants are cleaned up immediately

When an experiment concludes, the losing code path is removed — win or lose. Keeping dead variants in the codebase adds maintenance burden, makes the code harder to reason about, and turns temporary experiment flags into permanent conditions nobody understands.

06 // Flag Lifecycle & Debt

Flag debt compounds faster than code debt

Every active flag is a code path that must be maintained, tested (flag-on and flag-off), and reasoned about. Ten active release flags means your system has up to 1,024 possible behavioral states. Flag debt does not feel real until it is unmanageable.

1

Every flag has an expiry date at creation

A flag created without an agreed removal date will still be in the codebase two years later. At creation: set expected rollout completion date, name the owner of cleanup, and define what "done" looks like. This is non-negotiable for release and experiment flags.

2

Flag debt is tracked alongside code debt

Stale flags appear in sprint backlog and code review as explicitly as stale dependencies or TODO comments. The engineering team maintains a live inventory of all active flags: type, owner, age, and status. Flags older than their intended lifetime trigger a review.

3

The cleanup PR is part of the feature definition

A feature is done when the flag code path is removed, the config is deleted from the flag service, and the dead branch is cleaned up — not when it reaches 100% rollout. Cleanup is a tracked deliverable with the same weight as the feature itself.

07 // Observability of Flags

If you can't see what a flag is doing, you can't trust it

Without flag-level observability, debugging inconsistent behavior is nearly impossible. "Why is this user not seeing the feature?" becomes a multi-hour investigation through targeting rules, SDK versions, and cache states. Evaluations must leave a trail.

01
Log every evaluation with its targeting result
Every flag evaluation emits: flag name, resolved value, targeting rule that matched, and the user_id or tenant_id it was evaluated for. This is the only way to answer "why did user X get variant Y?" in under 30 seconds.
02
Track key metrics split by flag variant during rollout
For every flag controlling a user-facing behavior, the key business metric (error rate, conversion, latency) must be visible split by variant. This is what lets you catch a regression at 5% rollout before it reaches 50%. Without it, you are flying blind through the most dangerous window.
03
Flag evaluation must not add measurable latency
If your flag evaluation path adds more than 1–2ms, something is wrong — you are likely making a network call per evaluation. Evaluation is a local hash-map lookup against a cached config. Profile your SDK in production; the flag path should be invisible in your traces.
08 // Testing with Flags

A test suite that runs with one flag state is half a test suite

Every flag introduces a code branch. A test suite that only exercises the default flag state will pass CI and break production when the flag is enabled. Testing with flags is not optional — it is part of the correctness contract of the feature.

Test both flag states — on and off — in CI

Every test that exercises a flagged code path must run twice: once with the flag enabled, once disabled. Use flag overrides or test fixtures to set flag state explicitly. Never rely on default values in tests — defaults change.

Tests must not call the real flag service

Tests that resolve flags against a remote service are slow, flaky (network dependency), and couple test results to production flag configuration. Mock or override flag state locally in all test and CI environments. Flag evaluation in tests is a local operation.

Targeting rule changes go through code review

A targeting rule that accidentally enables a feature for all tenants instead of beta tenants is a production incident. Targeting rule changes have the same blast radius as code changes affecting user-facing behavior — they require the same review and the same rollout discipline.

The mental model

Deploy when engineering is ready.
Release when the business is ready.
Kill switches restore service without a rollback.
Flags accumulate debt. Schedule cleanup at creation.
A flag at 100% with no removal date is permanent technical debt.

Daniel Brasileiro