Skip to main content

DMARC aggregate reports (RUA): ingest and analysis

Publishing a DMARC record tells receivers what to do with mail that fails authentication. It does not tell you anything — until the receivers start reporting back. DMARC’s aggregate-report mechanism (RFC 7489 §7) is that feedback loop: every mailbox provider that received mail claiming to be from your domain mails you a daily XML rollup of what it saw, per sending source IP — how many messages, how SPF and DKIM evaluated, which disposition (none, quarantine, reject) it applied. Without those reports, a p=reject record is a blind bet. With them, you can answer the two questions a deliverability operation actually runs on: is my alignment holding in the wild, and who is sending as my domain that I never authorized. This page covers what RUA reports are, the transport shapes Orbit accepts, how the analysis endpoint turns them into a deliverability summary and policy recommendation, and where the result lands when you run a sender domain.

What a RUA report contains

A DMARC aggregate report is defined by RFC 7489 §7.2 and its Appendix C schema. Mailbox providers (Google, Microsoft, Yahoo, and the rest) batch a day’s worth of observations into one XML document per reporting domain, compress it with gzip, and mail it as an application/gzip attachment to every address listed in your _dmarc record’s rua= tag. Occasionally the attachment arrives as a .zip, and a hand-submitted or already-extracted report is just the raw XML. Each report carries three blocks:
  • Report metadata — the reporting provider, the report id, and the date range the rollup covers (usually one UTC day).
  • The policy the provider evaluated — the domain, the published p= value, and the pct sampling it applied. This is what receivers enforced, not what you meant to publish.
  • One record per source IP — message count, the header-from domain seen, the SPF and DKIM evaluation results, and the disposition the provider applied to that source’s mail.
A record passes DMARC when either its aligned DKIM or its aligned SPF passes (RFC 7489 §3.1). The aggregate row already carries that post-alignment verdict, so the analysis reads the provider’s conclusion rather than re-deriving it.

The three transport shapes Orbit accepts

POST /api/v1/email/dmarc/reports/analyze takes one or more reports in a single JSON body. Each report names its own transport encoding, because the bytes arrive in different shapes depending on how you collected them: utf8 is the default when you omit encoding; the same request can mix shapes, one encoding per report. The request body accepts up to 20 reports, each capped at 1 MB as submitted.

The 8 MB decompression ceiling

A gzip payload is attacker-supplied input with an effectively unbounded compression ratio — a small compressed stream can inflate to gigabytes of output (a classic decompression bomb). Orbit decodes gzip with a hard ceiling of 8 MB decompressed output per payload. The moment a payload would inflate past that bound, decompression fails; the report is counted as skipped, and the rest of the batch still analyzes. Eight megabytes sits comfortably above any legitimate DMARC aggregate — even a very high-volume sender’s daily report is a fraction of it — while keeping the worst-case allocation for a full 20-report batch bounded.

How the analysis pipeline works

The endpoint runs three steps, each tolerant of a single bad payload:
  1. Decode. Each report’s content is extracted to XML per its declared encoding. A report that fails to decode — bad base64, an over-cap gzip, an empty document — is counted and skipped. It never sinks the batch.
  2. Parse and summarize. The decoded XML documents are folded into one deliverability summary: total messages, DMARC pass rate, SPF-aligned and DKIM-aligned pass rates, disposition counts (none / quarantine / reject), the number of distinct source IPs seen, and the ranked list of sources sending DMARC-failing mail — each with its source IP, the header-from it claimed, its failing volume, and the dispositions receivers applied to it. A report whose XML is malformed at parse time joins the same skipped tally; the response reports an honest submittedReports / skippedReports split so you can see how much of what you sent was actually analyzed.
  3. Recommend the next policy step. The summary feeds the DMARC roll-out wizard: given your currently-published policy (none → quarantine → reject) and the observed pass rates, it returns advance, hold, or investigate, a confidence level, the specific blockers to fix first (a weak SPF-alignment rate, a weak DKIM-alignment rate, or a top failing source to identify), and a one-line next step. The wizard is deliberately conservative — it recommends advancing only when the pass rate is healthy over a meaningful sample, and it holds when the sample is too small to trust.
The endpoint is stateless: nothing is persisted. You POST reports, you get the analysis back, and you decide what to republish. topSpoofingSources (default 20, max 200) controls how many failing sources the ranking returns. Like the rest of the deliverability surface, the route requires an owner or admin role.

Where the analysis lands in the deliverability hub

The none → quarantine → reject progression is the one safe way to enforce DMARC, and the aggregate reports are the only input that makes each step defensible. The analysis feeds the same workflow the deliverability lab drives:
  • Starting at p=none. This is a monitoring posture — receivers report everything and block nothing. Read the DMARC pass rate first. Until it is healthy, advancing only turns a monitoring problem into a blocking problem.
  • Reading a failing source. Every entry in the ranked spoofing list is one of exactly two things: a legitimate sender you have not aligned (a marketing tool, a helpdesk, a payroll service sending with your header-from but outside your SPF and DKIM) or an actual spoofer. The first is onboarded — you either put it in SPF and DKIM on your domain or move it to a subdomain. The second is what enforcement exists for, and a healthy pass rate plus a tight spoofing list is the evidence that quarantine and then reject will only hurt the second kind.
  • Advancing to quarantine and reject. When the wizard reports advance at high confidence — pass rate at or above the healthy bar over a sufficient message sample — raise the policy. Re-run the same analysis after each republish; pct lets you sample the stricter policy on a fraction of failing mail before applying it to all of it.
  • When a domain is attacked or downgraded. Read the ranked failing sources first, then the disposition counts. A spike of none-disposition failures at an unfamiliar source IP means the spoofing is flowing unimpeded — your published policy is too weak or your alignment data is too young to move. A block of failures already at quarantine or reject means enforcement is doing its job, and the remaining question is whether any of those sources are legitimate senders to onboard.
Two sibling pages frame this in the wider picture: Sender warming and reputation covers the engagement-side signals a warming sender is graded on, and Email transport security covers the connection-plane counterpart (TLS-RPT) whose aggregate ingestion mirrors this one.

Publishing the rua= record that routes reports to you

Mailbox providers learn where to send reports from the _dmarc TXT record on your domain — the same record that publishes your policy:
  • rua= is a comma-separated list of mailto: addresses (an https: endpoint is also valid per the RFC). Each one receives a copy of every provider’s aggregate.
  • The address must accept report attachments — gzip’d XML, daily, from many providers. Point it at a mailbox or collector you monitor for infrastructure reports; many tenants run a dedicated address for exactly this.
  • Keep rua= in the record from the first day at p=none. Reports only begin after receivers see the tag, so a domain that publishes p=quarantine without rua= has enforcement and no visibility into what it is enforcing.
Publish one record, never two — a duplicate _dmarc TXT makes receivers ignore both. The deliverability DNS check grades the published record; once reports start arriving, POST them to the endpoint above and read the summary.

Worked example: first reports on a new sender domain

1. Publish with reporting from day one. Put your sender domain at p=none with a rua= address you monitor. Leave it there — nothing is being blocked, and receivers start mailing aggregates within about a day. 2. Collect a week of aggregates. Pull the gzip attachments out of the report mailbox, or have your collector forward the decoded XML. 3. Submit the batch.
4. Read the summary, then the wizard. Confirm submittedReports matches what you sent and skippedReports is low. Check the DMARC pass rate and the spoofing ranking: every failing source is either a sender to align or a spoofer to enforce against. The recommendation tells you whether the data supports moving to quarantine yet — and, when it does not, names which mechanism (SPF alignment, DKIM alignment, or a specific failing source) is blocking the move. 5. Advance, republish, repeat. Move p=none → p=quarantine; pct=25 when the wizard recommends advancing, raise pct as the pass rate holds, and finish at p=reject. Each republish re-arms the loop: new aggregates confirm the stricter policy is only catching mail that deserves it.

See also