Using Weighted DNS for Gradual Traffic Migration

Problem Statement

A hard DNS flip sends 100% of traffic to a brand-new origin in one move, so any latent capacity, configuration, or data-layer issue hits every user at once. Weighted DNS lets you publish two records for the same hostname and split resolution by a configurable ratio, so you can send 5%, then 25%, then 100% of traffic to the new origin while watching real metrics. This page, part of Zero-Downtime Cutover Plans, shows how to run that gradual, canary-style cutover with Route 53 weighted routing or an equivalent provider.

Weighted DNS traffic split A weighted record set splits resolution between the old and new origins by ratio, ramping the new origin from a small canary to full traffic. Weighted DNS Traffic Split Weighted record Old origin New origin example.com A weight 95 weight 5 (canary) Resolvers Ramp new-origin weight up as metrics stay healthy; drop to 0 to roll back
Resolution is split by weight; the new origin starts as a small canary and its weight is ramped up as metrics hold healthy.

When to Use This Approach

  • The new origin is functionally identical to the old one for the same hostname and you can run both in parallel.
  • You want to limit blast radius by exposing a small percentage of real traffic first.
  • You need a rollback that is a single weight change rather than a re-publish of the original record.
  • Your provider supports weighted routing (Route 53 weighted records, or equivalents on Cloudflare load balancing, NS1, or Azure Traffic Manager).
  • You have lowered TTL ahead of the ramp (see How to Lower DNS TTL Before Migration) so weight changes take effect quickly.

Step-by-Step Instructions

1. Lower TTL So Weight Changes Take Effect Fast

Weighted routing only ramps as quickly as resolvers re-query, which is governed by TTL. Set a short TTL (60–120 seconds) on both weighted records before you start so each weight adjustment propagates within a couple of minutes.

# Confirm the short TTL is live before ramping
dig @1.1.1.1 example.com A +noall +answer
# example.com.   60   IN   A   198.51.100.20

2. Create the Two Weighted Records

Publish two records with the same name and type but different SetIdentifier values and weights. Start with the new origin at a small weight so it receives only a canary slice of traffic.

# Old origin — weight 95
aws route53 change-resource-record-sets --hosted-zone-id Z123EXAMPLE --change-batch '{
  "Changes": [{ "Action": "UPSERT", "ResourceRecordSet": {
    "Name": "example.com.", "Type": "A", "TTL": 60,
    "SetIdentifier": "old-origin", "Weight": 95,
    "ResourceRecords": [{"Value": "198.51.100.20"}] }}]}'

# New origin — weight 5 (the canary)
aws route53 change-resource-record-sets --hosted-zone-id Z123EXAMPLE --change-batch '{
  "Changes": [{ "Action": "UPSERT", "ResourceRecordSet": {
    "Name": "example.com.", "Type": "A", "TTL": 60,
    "SetIdentifier": "new-origin", "Weight": 5,
    "ResourceRecords": [{"Value": "203.0.113.10"}] }}]}'

3. Validate the Canary Slice

With ~5% of resolutions hitting the new origin, watch new-origin error rates, latency, and application logs against the old-origin baseline. Hold at this weight long enough to cover real traffic patterns before ramping further.

# Sample resolution repeatedly to confirm the ~95/5 split is observable
for i in $(seq 1 20); do dig @8.8.8.8 example.com A +short; done | sort | uniq -c

4. Ramp the Weight Upward in Stages

If the canary holds healthy, increase the new-origin weight in stages — for example 5 → 25 → 50 → 100 — pausing at each stage to validate. Adjust both records each time so the ratio is explicit.

# Stage to 25% new origin: old weight 75, new weight 25
aws route53 change-resource-record-sets --hosted-zone-id Z123EXAMPLE --change-batch '{
  "Changes": [
    { "Action": "UPSERT", "ResourceRecordSet": { "Name": "example.com.", "Type": "A",
      "TTL": 60, "SetIdentifier": "old-origin", "Weight": 75,
      "ResourceRecords": [{"Value": "198.51.100.20"}] }},
    { "Action": "UPSERT", "ResourceRecordSet": { "Name": "example.com.", "Type": "A",
      "TTL": 60, "SetIdentifier": "new-origin", "Weight": 25,
      "ResourceRecords": [{"Value": "203.0.113.10"}] }}
  ]}'

5. Complete or Roll Back the Cutover

To finish, set the old-origin weight to 0 and the new-origin weight to a positive value so all traffic resolves to the new origin. To roll back at any stage, do the inverse — drop the new-origin weight to 0. A weighted ramp is a per-host strategy and pairs well with the environment-swap pattern in Implementing Blue-Green Deployments for Site Migrations.

# Finish: old origin to 0, new origin carries everything
aws route53 change-resource-record-sets --hosted-zone-id Z123EXAMPLE --change-batch '{
  "Changes": [
    { "Action": "UPSERT", "ResourceRecordSet": { "Name": "example.com.", "Type": "A",
      "TTL": 60, "SetIdentifier": "old-origin", "Weight": 0,
      "ResourceRecords": [{"Value": "198.51.100.20"}] }},
    { "Action": "UPSERT", "ResourceRecordSet": { "Name": "example.com.", "Type": "A",
      "TTL": 60, "SetIdentifier": "new-origin", "Weight": 100,
      "ResourceRecords": [{"Value": "203.0.113.10"}] }}
  ]}'

The ramp is only as useful as the gate between stages, and the gate has to be a number rather than a feeling. Each stage should hold until it has produced a sample large enough to distinguish a real regression from noise — which at 5% of traffic takes considerably longer than at 50%.

Ramp stages with the evidence required to advance each one Four ramp stages at five, twenty-five, fifty and one hundred percent, each showing the share of traffic on the new origin, how long it takes that stage to gather a meaningful sample, and the condition that must hold before advancing. Advance on evidence, not on the clock 5% 25% 50% 100% canary slice real load appears capacity honestly tested old weight = 0 needs the longest soak watch p95, not just errors connection pools bite here keep old origin warm ≥ 2 000 requests error rate ≤ baseline p95 within tolerance all stages green Rollback at any stage is one weight change to zero — which is why each stage is cheap to hold and expensive to skip.
A 5% stage advanced after ten minutes has usually seen too little traffic to prove anything; the smallest slice needs the longest wait.

Worked Example

A retailer migrates example.com from 198.51.100.20 to a new platform at 203.0.113.10. TTL is pre-lowered to 60s.

Day 1, canary at weight 5. Sampling 8.8.8.8 twenty times shows the split is live:

$ for i in $(seq 1 20); do dig @8.8.8.8 example.com A +short; done | sort | uniq -c
     19 198.51.100.20
      1 203.0.113.10

New-origin error rate sits at 0.1%, matching the old origin, so the team ramps to 25%, then 50% over the next two days, validating at each stage. On day 4 they cut old-origin weight to 0:

$ for i in $(seq 1 20); do dig @8.8.8.8 example.com A +short; done | sort | uniq -c
     20 203.0.113.10

All sampled resolutions now return the new origin. Had error rate spiked at any stage, setting the new-origin weight to 0 would have reverted the slice within one TTL window.

One property of weighted DNS deserves stating plainly before you rely on it for a user-facing experiment: the weight applies to resolvers, not to people. A resolver that draws the new origin caches that answer and sends everyone behind it to the new origin for the life of the entry, so the split you observe in aggregate is not a random sample of users.

How a weighted split maps onto real users A five percent weight resolved by three resolvers: the one that happens to draw the new origin sends every user behind it there, so a large corporate resolver can put far more than five percent of people on the new origin. 5% of resolvers is not 5% of people Consumer ISP resolver drew old origin — 40k users Public resolver drew old origin — 35k users Large corporate resolver drew NEW origin — 25k users 75% of users still on the legacy origin 25% on the new one The weight was set to 5%. One resolver drawing the new record put a quarter of the audience behind it — and they are correlated, all from the same company, on the same network, in the same timezone. Measure the slice from origin request counts, never from the configured weight.
Because a resolver's draw is sticky for the life of its cache, the users on each origin are a clustered sample rather than a random one — fine for a capacity canary, unreliable as an A/B test.

Verification

# Sampled distribution should track the configured weights
for i in $(seq 1 50); do dig @8.8.8.8 example.com A +short; done | sort | uniq -c

# Confirm the weighted set itself is published correctly at the provider
aws route53 list-resource-record-sets --hosted-zone-id Z123EXAMPLE \
  --query "ResourceRecordSets[?Name=='example.com.']"

Each ramp stage passes when the sampled resolution ratio approximates the configured weights and new-origin error rate and latency stay within your thresholds; the cutover is complete when 100% of samples return the new origin and the old-origin weight is 0.

FAQ

Why don’t I see exactly the weight ratio in my dig samples? Weighted routing is probabilistic per query and influenced by resolver caching, so a small sample drifts from the configured ratio. Sample 50–100 times across multiple resolvers and the observed distribution will converge on the weights; very short windows are noisy by nature.

How fast can I roll back a weighted cutover? Rollback is one change: set the new-origin weight to 0. Because you lowered TTL before starting, resolvers re-query within that TTL window — typically one to two minutes — so a bad stage drains far faster than re-publishing an original A record from scratch.

Does weighted DNS work for a hostname behind a CDN? Usually not in the way you expect, because the resolver caches the CDN’s edge address rather than your origin, so your weights never enter the picture for end users. Where the CDN proxies the hostname, run the split at the CDN’s origin configuration or load-balancing layer instead — the staged ramp and the evidence gates are identical, only the control surface moves. Weighted DNS remains the right tool for hostnames resolved directly, including API endpoints and any unproxied subdomain.

Can I use weighted DNS without Route 53? Yes. Cloudflare load balancing, NS1, and Azure Traffic Manager all offer weighted or proportional routing using the same principle of two records ramped by ratio. The commands differ but the staged 5 → 25 → 50 → 100 ramp and the lowered-TTL prerequisite are identical.

What should I watch during the canary that I would not watch normally? Watch the things that only appear under real traffic and only on a new host: connection-pool saturation, TLS handshake failures from clients your test suite does not represent, and any log line mentioning a hostname or path that should no longer exist. Standard error-rate dashboards catch the loud failures, but the canary exists to surface the quiet ones — a payment provider rejecting calls from an unfamiliar source IP, a rate limiter counting the new origin as a single abusive client, a scheduled job that has now started running on both origins at once. Read the new origin’s logs directly during the first stage rather than relying on aggregated metrics.

How do I stop a weighted set from silently persisting after the migration? Delete the zero-weight record rather than leaving it at zero. A record with weight 0 is inert but still published, and the next person to look at the zone has no way to know whether it is a deliberate standby or forgotten debris. Worse, some providers treat an all-zero weighted set as an instruction to return every record, so if someone later zeroes the surviving entry while tidying up, traffic silently resumes flowing to a decommissioned address. Remove the old SetIdentifier entirely once the origin is retired, and record the deletion alongside the cutover.

Related

← Back to Zero-Downtime Cutover Plans