WarpLink
Trust and reliability

Architecture and Failure Modes

How WarpLink serves redirects from an edge cache, records clicks through a queue, and what keeps working when the database, API, or click queue is down.

Cached links keep redirecting when the database is down, and keep redirecting when the dashboard and API are down. A link that is not in the cache needs the database, so during a database outage it answers 503 with a Retry-After header instead of a false "not found". This page describes the design and each failure case in plain terms.

Two paths

WarpLink runs two separate paths, so a problem in one does not take down the other.

PathWhat it servesWhat it depends on
Redirect pathShort-link taps from visitorsThe edge key-value cache. The database only on a cache miss.
API and dashboard pathThe dashboard, the REST API, the MCP server, and the SDK calls for link resolution and attributionThe database, and the same edge network

How a redirect works

  1. The visitor taps a short link. The nearest edge location receives the request.
  2. The redirect service reads the link's configuration from the edge key-value cache. For a custom domain it first reads the domain's serving status from the same cache.
  3. It screens the user agent, chooses the destination for the visitor's operating system, and returns a 302.
  4. After the response is on its way, it puts one click event on a queue and, for mobile visitors, stores the deferred-link payload in the cache. Neither step makes the visitor wait.

Every link is written to the database first and then to the cache. If a cache write fails, it is retried once and then placed on a repair queue. A repair queue message re-attempts the write later.

Click recording

A separate consumer reads click events from the queue in batches of up to 100 and writes them to the database. Click uniqueness is decided in the database at write time. Each event carries the time of the click, so a late write still lands in the correct billing period and the correct day of your analytics.

When a dependency is down

What is downCached linksUncached linksClick recordingSDK link resolution and attribution
DatabaseKeep redirecting503 with Retry-After: 5, not 404Events wait on the queue and are retriedFail until the database returns
Dashboard and APIKeep redirectingKeep redirecting from the cache or databaseKeep recordingFail until the API returns
Click queue delayedKeep redirectingKeep redirectingAnalytics arrive late, not lostUnaffected
Edge cacheRedirect service returns an errorRedirect service returns an errorNot reachedUnaffected

The cache is the one dependency on the redirect path itself. If the cache cannot be read, the redirect service has no link to serve and answers with an error.

Database down

A link already in the cache redirects exactly as usual, because the redirect path never reads the database for it. A link that is not in the cache needs one database read. That read has a 2-second limit per attempt and two attempts. If both fail, the visitor gets a 503 with Retry-After: 5. We do not answer 404 in that case. A 404 tells the visitor, and every search crawler behind them, that a live link is dead, and that answer gets cached.

Click events keep flowing onto the queue during a database outage. The consumer cannot write them, so it hands the batch back to the queue with a growing delay: 15 seconds, doubling, up to 5 minutes per wait. A batch gets five retries, which spans roughly eight minutes. After that the batch moves to a second queue, which retries each click separately so one bad event cannot block the rest. A click that still cannot be written moves to a holding queue that has no automatic consumer. Nothing there is deleted. An operator drains it after the database is back.

Dashboard and API down

Redirects do not use the dashboard or the API, so they are unaffected.

Everything an app asks over the API is affected, because the SDK calls the API:

  • Opening a Universal Link or App Link. The operating system opens the app, and the app then calls GET /v1/links/resolve/{slug} to learn the deep link and its parameters. That call needs the API. Without it, the SDK cannot resolve the link.
  • Install attribution (POST /v1/attribution/match).
  • Key validation at configure time.

On failure, the iOS and Android SDKs make up to three attempts: a first attempt with a 4-second limit, then two retries with a 3-second limit each, separated by waits of about 0.5 and 1 second. The whole run stops at 12 seconds. Only network errors and 5xx responses are retried in the current release. Other 4xx responses are refusals that would repeat, so they are not retried. Retrying a 429 is covered under request limits below. When the run ends in failure, the SDK delivers the error to your onLink callback and your app decides what to show. The React Native and Flutter SDKs use the native SDKs for this. An attribution check that produced no usable answer is not marked complete, so the SDK runs it again on the next app launch. See Deferred deep links.

Click queue delayed

Visitors are unaffected. Analytics are late and catch up when the queue drains. The retry and holding-queue behavior is the same as in the database case above.

One case loses a click. If the queue itself cannot accept the event at the moment of the tap, that one click is not recorded. The redirect still completes.

Edge cache unavailable

The redirect service reads every link from the cache. If a read fails, the request ends in an error response, and no database fallback runs.

Request limits for SDK calls

An SDK key can call only the three SDK endpoints: key validation, link resolution, and attribution match. Every other route refuses it.

SDK calls are limited to 60 requests per minute per key per client IP address. Each device has its own address, so the limit applies per device, not across your whole install base. Devices behind one public IP address, such as an office or a household network, share that allowance. API keys, which servers and CI jobs use, are limited to 60 requests per minute per key. See Rate limiting.

From the next SDK release, an SDK that receives a 429 response retries once, waits for the time in the Retry-After header, and waits no more than 5 seconds.

Health checks

The redirect service exposes /health for liveness and /health?deep=1 for readiness. The deep check probes the database and the cache, each with its own deadline.

What is not covered here

This page describes the design and the failure behavior in the code. It does not state an uptime figure. We do not publish one until we can publish a measured history. For measured redirect timings, see Performance and methodology.

On this page