All systems operational.
A public look at the health of every Marqee surface — the marketing web, the member mobile app, Backstage where strategists work, the API underneath, and the concierge chat that ties it together. Checked continuously and updated in real time. Last refresh a moment ago.
Every surface, right now.
A ninety-day view of each Marqee system. Each dot is one day; green means the day closed operational, amber means a partial-impact window, red means an outage.
Where we're aiming.
These are the targets we hold ourselves to across every surface — the same numbers you'd expect from an infrastructure partner you're trusting with real work.
These are the targets we're operating to in Marqee's first year. Real, calculated numbers will publish here once ninety days of continuous measurement have elapsed — no rounded-up hype, no invented averages, just what actually happened. The measurement is simple: total minutes in the window, minus the minutes any surface was down or degraded, divided by total minutes. If a surface was only partially affected, we count the impacted minutes at their real weight, not zero.
Watched from the inside and the outside.
Every surface above is probed continuously from three independent regions, plus a fourth internal probe that behaves like a real member — signing in, opening a page, sending a chat message, then signing out. If any probe fails twice in a row, an on-call engineer is paged within sixty seconds. The board updates the moment a probe changes state, so what you see here is what our own team is seeing on their screens.
The record, unedited.
Three minor incidents in the last ninety days, all resolved. We write these up because trust is built one small event at a time, and the post-mortem is what turns an outage into a lesson. Every entry below shows the impact window, the affected surface, what actually broke, and what we changed so the same failure mode can't repeat. If an incident affected you and you're not sure whether it's the one described, reach out and a strategist will confirm.
A queue backup in the message-delivery layer caused new messages between members and strategists to appear in the recipient's inbox with a delay of 30 to 90 seconds. No messages were dropped or lost, and file attachments continued to upload normally. The backlog cleared once the affected worker pool was scaled and the stuck consumer was recycled; end-to-end delivery returned to under two seconds. Root cause was a single consumer whose lease had expired without releasing its partition, a case our health-check hadn't been configured to catch. We've since raised the queue-depth alert threshold, replaced the lease heuristic with an explicit heartbeat, and added a synthetic round-trip check that pages an on-call engineer within 60 seconds of any delivery lag over five seconds.
During a routine CDN edge migration, roughly one in six requests to the marketing homepage returned in 4 to 7 seconds instead of under a second. Interior pages and the member web app were unaffected because they route through a separate origin, so signed-in members did not see the slowdown. We paused the migration, rolled the affected edge region back to the previous configuration, and completed the migration overnight in a maintenance window without further impact. A phased rollout gate has been added to the migration playbook so no more than 5% of edge traffic can shift in a single window going forward, and every migration now runs against a synthetic homepage load test before promotion.
A session-store deploy left a small percentage of strategists unable to sign in on their first attempt; a retry succeeded on the second attempt every time. The deploy was rolled back within the window, and the underlying cache-invalidation ordering issue — where a freshly-issued session token could race the previous one out of the cache — was patched the same evening. No member-facing surface was affected, and no session data was corrupted. We've added a canary sign-in probe that runs against every deploy of the session store before it takes traffic, and moved the affected write path to a serialised queue so the ordering can't repeat.
Nothing scheduled right now.
When a window is needed, we take it late on a Sunday and give you advance notice on this page and by email.
Standard window: Sundays 04:00–06:00 UTC
If we need to take a surface offline for a deploy, migration, or infrastructure upgrade that can't be rolled forward, we do it late on a Sunday in the quietest window across our member timezones. You'll see the window listed here at least 48 hours in advance, and members with active searches receive an email the day before with the exact start time in their local zone and an estimate of how long each affected surface is expected to be offline.
Get notified when something changes.
Pick email if you want a note in your inbox, or the RSS feed if you already have a reader wired into your workflow.
Email me when there's an incident
One email at the start of an incident, one when it's resolved, one when the post-mortem is published. No marketing, ever.
Three states, plainly defined.
So a green pill above always means the same thing, and a yellow one is never a shrug. If the definition here doesn't match what you're seeing, that's a bug worth telling us about — send a note and a strategist will loop in engineering. We'd rather over-report a rough patch than dress it up.
Every key member flow is working end to end — signing in, opening your dashboard, chatting with your strategist, uploading a file, submitting a change. This is the default and where we want to stay.
Something is partially impacted — a page is slower than it should be, one region is affected, or a non-critical feature is off. The core member flows still work, and we're already actively fixing it.
The surface is unavailable and members can't use it right now. This is a serious event: we'll post an update on this page within minutes, keep updating as we go, and publish a full post-mortem within 48 hours.