Architecture and Failure Modes
How WarpLink serves redirects from an edge cache, records clicks through a queue, and what keeps working when the database, API, or click queue is down.
Cached links keep redirecting when the database is down, and keep redirecting when the dashboard and API are down. A link that is not in the cache needs the database, so during a database outage it answers 503 with a Retry-After header instead of a false "not found". This page describes the design and each failure case in plain terms.
Two paths
WarpLink runs two separate paths, so a problem in one does not take down the other.
| Path | What it serves | What it depends on |
|---|---|---|
| Redirect path | Short-link taps from visitors | The edge key-value cache. The database only on a cache miss. |
| API and dashboard path | The dashboard, the REST API, the MCP server, and the SDK calls for link resolution and attribution | The database, and the same edge network |
How a redirect works
- The visitor taps a short link. The nearest edge location receives the request.
- The redirect service reads the link's configuration from the edge key-value cache. For a custom domain it first reads the domain's serving status from the same cache.
- It screens the user agent, chooses the destination for the visitor's operating system, and returns a 302.
- After the response is on its way, it puts one click event on a queue and, for mobile visitors, stores the deferred-link payload in the cache. Neither step makes the visitor wait.
Every link is written to the database first and then to the cache. If a cache write fails, it is retried once and then placed on a repair queue. A repair queue message re-attempts the write later.
Click recording
A separate consumer reads click events from the queue in batches of up to 100 and writes them to the database. Click uniqueness is decided in the database at write time. Each event carries the time of the click, so a late write still lands in the correct billing period and the correct day of your analytics.
When a dependency is down
| What is down | Cached links | Uncached links | Click recording | SDK link resolution and attribution |
|---|---|---|---|---|
| Database | Keep redirecting | 503 with Retry-After: 5, not 404 | Events wait on the queue and are retried | Fail until the database returns |
| Dashboard and API | Keep redirecting | Keep redirecting from the cache or database | Keep recording | Fail until the API returns |
| Click queue delayed | Keep redirecting | Keep redirecting | Analytics arrive late, not lost | Unaffected |
| Edge cache | Redirect service returns an error | Redirect service returns an error | Not reached | Unaffected |
The cache is the one dependency on the redirect path itself. If the cache cannot be read, the redirect service has no link to serve and answers with an error.
Database down
A link already in the cache redirects exactly as usual, because the redirect path never reads the database for it. A link that is not in the cache needs one database read. That read has a 2-second limit per attempt and two attempts. If both fail, the visitor gets a 503 with Retry-After: 5. We do not answer 404 in that case. A 404 tells the visitor, and every search crawler behind them, that a live link is dead, and that answer gets cached.
Click events keep flowing onto the queue during a database outage. The consumer cannot write them, so it hands the batch back to the queue with a growing delay: 15 seconds, doubling, up to 5 minutes per wait. A batch gets five retries, which spans roughly eight minutes. After that the batch moves to a second queue, which retries each click separately so one bad event cannot block the rest. A click that still cannot be written moves to a holding queue that has no automatic consumer. Nothing there is deleted. An operator drains it after the database is back.
Dashboard and API down
Redirects do not use the dashboard or the API, so they are unaffected.
Everything an app asks over the API is affected, because the SDK calls the API:
- Opening a Universal Link or App Link. The operating system opens the app, and the app then calls
GET /v1/links/resolve/{slug}to learn the deep link and its parameters. That call needs the API. Without it, the SDK cannot resolve the link. - Install attribution (
POST /v1/attribution/match). - Key validation at configure time.
On failure, the iOS and Android SDKs make up to three attempts: a first attempt with a 4-second limit, then two retries with a 3-second limit each, separated by waits of about 0.5 and 1 second. The whole run stops at 12 seconds. Only network errors and 5xx responses are retried in the current release. Other 4xx responses are refusals that would repeat, so they are not retried. Retrying a 429 is covered under request limits below. When the run ends in failure, the SDK delivers the error to your onLink callback and your app decides what to show. The React Native and Flutter SDKs use the native SDKs for this. An attribution check that produced no usable answer is not marked complete, so the SDK runs it again on the next app launch. See Deferred deep links.
Click queue delayed
Visitors are unaffected. Analytics are late and catch up when the queue drains. The retry and holding-queue behavior is the same as in the database case above.
One case loses a click. If the queue itself cannot accept the event at the moment of the tap, that one click is not recorded. The redirect still completes.
Edge cache unavailable
The redirect service reads every link from the cache. If a read fails, the request ends in an error response, and no database fallback runs.
Request limits for SDK calls
An SDK key can call only the three SDK endpoints: key validation, link resolution, and attribution match. Every other route refuses it.
SDK calls are limited to 60 requests per minute per key per client IP address. Each device has its own address, so the limit applies per device, not across your whole install base. Devices behind one public IP address, such as an office or a household network, share that allowance. API keys, which servers and CI jobs use, are limited to 60 requests per minute per key. See Rate limiting.
From the next SDK release, an SDK that receives a 429 response retries once, waits for the time in the Retry-After header, and waits no more than 5 seconds.
Health checks
The redirect service exposes /health for liveness and /health?deep=1 for readiness. The deep check probes the database and the cache, each with its own deadline.
What is not covered here
This page describes the design and the failure behavior in the code. It does not state an uptime figure. We do not publish one until we can publish a measured history. For measured redirect timings, see Performance and methodology.
Performance and Methodology
What WarpLink's sub-10ms redirect figure measures, a dated warm-link test with its sample size, client latency from one location, and measured SDK sizes.
Attribution Accuracy
Which install matches are deterministic, the known failure modes, and the method and device checklist for measuring correct, wrong, and unmatched installs.