I Tested an LLM on Our DMARC Data. Here’s What I Found
Part of Our Own Estate, where we publish what DMARC reports say about the domains we run.
I connected Claude to the DMARKOFF MCP server and pointed it at one of our own domains: a newsletter subdomain sending through Amazon SES, 22,991 messages in the two weeks to 9 September, 53 distinct sending sources, p=none. I know every source on that domain, which makes it a fair test. I can check each conclusion the model reaches against what is actually happening.
I can hear the objection: people from GlockApps and DMARKOFF, and the policy is still none. It is a new subdomain, and we ran it the way we tell everyone else to run one: publish p=none, collect reports, account for every legitimate sender, then move. The less flattering part is that collecting has taken five months, which is longer than it needed to. That is one of the questions this article ends on, and it is there because the data put it there, not because it makes a tidy ending.
What an LLM Can and Cannot Conclude from DMARC Aggregate Data

The second run. The prompt asks for a headline, then for the field behind every claim, sorted by whether it was read from DNS, read from reports, or inferred. The headline is the model's own conclusion: on this domain, the Critical severity is a volume anomaly, not a configuration defect.
The result was not "the model is right" or "the model is wrong". It was more specific than that. On the first pass, the model produced six statements that a person who runs DMARC would not make, and every one of them was corrected by a field the tools had already returned, usually in the same response. None of the corrections needed information from outside the reports. That is worth writing down, because it says something useful about where the boundary actually sits, and about what a DMARC tool has to hand a model if the model is going to stay on the right side of it.
What Is in an Aggregate Report
An RUA report is a receiver telling you, per source IP and per day, how many messages it saw carrying your domain in the From header, what SPF and DKIM returned, which domains and selectors were involved, whether either result aligned, and what it did with the message. That is the whole content. No subject lines, no recipients, no message bodies, no indication of whether the mail was wanted, and nothing about who operates a sending IP beyond what you can look up yourself.
So every statement about a DMARC report belongs to one of three kinds: something you read directly from the report or from DNS, something you infer with a stated confidence, and something you cannot settle without a person or a second data source. A model that keeps those apart is useful. A model that blends them is a liability, because it writes all three with the same fluency.
How the Tools Are Shaped Around That Split
The MCP server is built on the assumption that a model will try to blend them, so the tools are designed to make the seams visible.
They return evaluations, not verdicts. get_domain_source_details gives the SPF result with its return path and scope, the DKIM result with its signing domain and selector, the disposition, and the receiver's own policy override comment. No field says whether a source is malicious. The model has to reason from what a person would reason from, and the fields are there to be quoted back.
Record health and activity health are separate objects. A DNS finding such as a lookup count or a malformed record is deterministic and comes with the offending value. An activity finding such as a compliance drop is statistical and comes with the quantile it was measured against. Those deserve different confidence, so they are not mixed into one score.
Findings carry the context that qualifies them. get_domain_full_data returns previousPeriod next to the current numbers, so "is this new" is answered in the same object as "is this bad". Every anomaly carries lastReportDate, so "compliance dropped" can be weighed against "the report has not arrived". get_smtp_rejections says in its own description that an empty result means no receiver reported a rejection, not that nothing bounced.
Classification is a label with its evidence attached. Sources are grouped into known, unknown, and forward, and each row also carries sourceReason in plain words along with the full authentication detail. When the label and the evidence disagree, the evidence is in the same row, which is what lets a careful reader, or a careful prompt, overrule it.
The workflow in the server instructions orders the questions: overview, domains by severity, one domain in full, the timeline, the anomaly comparison, then the drill-downs and the symptom tools. The order exists so that a source-level conclusion is not reached before the record-level context that explains it.
Every correction below came out of one of those decisions.
What the Model Reads Correctly
Forwarded mail. The domain had 1,509 forwarded messages across 48 forwarding sources, from Google and Microsoft down to single messages through Apple, Fastmail, Zoho, and a dozen small hosts. Nearly all of them: SPF fail, DKIM pass, DKIM aligned, DMARC pass. The model read this correctly without prompting. A forwarder re-sends the message from its own IP, which the return-path domain does not authorise, so SPF fails. The body and the signed headers arrive intact, so DKIM verifies and DMARC passes on the DKIM identifier. It did not raise an alarm about the 2,910 SPF failures in the summary, which is the number a dashboard shows first and the number a hurried reader reacts to.
An SPF record that needs nothing. get_spf_usage resolved the record live and laid the traffic onto it:
"record": "v=spf1 include:amazonses.com ~all", "lookupCount": 1
include:amazonses.com → ip4:23.251.224.0/19 total 20088, spfPass 20081
Thirteen ip4 ranges inside the include, one of them carrying every message. The model said the record is at one lookup out of ten, has no unused mechanisms worth removing, and needs no work. That is the right answer, and the interesting part is that it is a negative finding: the tool exists to clean up bloated records, and the model correctly declined to invent a cleanup.
An empty result read as an empty result. get_smtp_rejections returned nothing. The model reported that no receiver had attached a "Sender requirement failed" comment in the period, and added that this is not evidence that nothing bounced. That distinction is in the tool's description, and the model used it rather than converting silence into good news.
A volume spike that is not an incident. This is the one I did not expect to work as cleanly as it did.
messages 22991, compliancePct 87.4
previousPeriod: messages 5100, compliancePct 87.24
changes: messagesPct 350.8, compliancePctPoints 0.16
Volume up 351 percent, compliance flat to within a fifth of a point. The model's conclusion: a campaign went out, the authentication picture did not change, nothing to investigate. Before previousPeriod existed, the same model would have had to ask for a second period and compare by hand, and in my experience that is exactly the step it skips.
Where the First Pass Goes Wrong
"Twelve percent of your mail is failing DMARC"
The summary returns two percentages and a failure count:
"messages": 22991, "compliancePct": 87.4,
"pctEligibleForPolicy": 93.53,
"dmarcFail": 2898
The model opened with 12.6 percent of mail failing DMARC. Check it against the daily timeline: the dmarcFail column sums to 1,389 across the period, and forwarded sums to 1,509. Together they are 2,898. The summary's dmarcFail is messages - compliant, and compliant deliberately excludes forwarded mail, so forwarded messages that passed DMARC on their DKIM signature are counted in the summary's failure number. Genuine failures are 1,389, which is 6.0 percent. pctEligibleForPolicy, 93.53, is the number that answers "what would happen if I enforced".
Nothing here is wrong or hidden; both definitions are documented, and each is the right number for a different question. The failure is that "compliance" has no fixed meaning across DMARC vendors, so a model trained on all of their documentation reaches for whichever definition its sentence needs, and states it without saying which one it took. A number without its denominator is the most common way a model is wrong while every digit is correct. The fix is a prompt that asks which field a percentage came from, and the fields are all in one response.
"The policy is doing its job"
Every disposition in the data is none. Failing mail from unrecognised sources, gateway-mangled copies, all of it delivered. The model's draft: mail is being delivered normally, no enforcement problems.
The domain publishes p=none. There is no enforcement to have a problem with. get_record_history returned a single snapshot, dated when monitoring started in April, with no changes since:
"date": "2026-04-20", "policy": "none", "sp": "none", "pct": 100, "fo": 1
One snapshot with no changedFromPrevious entries means the record has been the same for five months. The model had read the record; it did not connect the record to the dispositions. Under p=none, a disposition: none tells you nothing about protection, and neither does a clean-looking delivery picture. This one matters because it is the direction of error that costs money: a confident all-clear on a domain that is not defended.
"You are being spoofed through Google"
Google appears twice in the source list, in two different classifications, at almost the same volume. A source can sit in more than one tab, which is why the tab counts add up to 59 against 53 distinct sources:

Nine unknown sources, 1,389 messages. Read as a list of names, this looks like a spoofing problem. All but eleven of those messages are our own mail coming back through gateways and forwarders.

The same period, the forward tab. Google appears here too, at almost identical volume, with DMARC passing on all but two of 1,249 messages. SPF fails in both tables. The DKIM column is what separates them.
The model's first reading of the unknown table was unauthorised mail sent through Google infrastructure. The drill-down shows two different things inside it, and the rows say which is which.

Source details, filtered to unknown. One From domain, two different stories. The top rows carry our own Return-Path and our own SES selector, with the signature failing: our mail, modified in transit. The bottom rows carry a subscriber's employer domain, where SPF and DKIM both pass for that domain and neither aligns with the From.
Some carry the domain's own return path and its real SES selector, with the signature failing:
"sourceReason": "relay_broke_dkim",
"spfAuth": [{ "returnPath": "bounce.news.glockapps.co", "result": "softfail" }],
"dkimAuth": [{ "domain": "news.glockapps.co", "selector": "dkayeghju235...", "result": "fail" },
{ "domain": "amazonses.com", "selector": "kra23psoka5...", "result": "fail" }]
A third party does not attach a signature naming your real selector; it has no key and no reason. A message carrying your selector and failing verification is your own message, signed correctly at origin and modified in transit. The same shape shows up on Check Point Harmony, INKY, Perception Point, Sophos, and Barracuda rows: inline security gateways that rewrite links, insert banners, and re-emit the message, breaking the body hash on the way through.
The rest have a different shape entirely:
"spfAuth": [{ "returnPath": "<a CRM vendor's domain>", "headerFrom": "news.glockapps.co", "result": "pass" }],
"dkimAuth": [{ "domain": "<the same vendor's domain>", "selector": "google", "result": "pass" }]
SPF passes, DKIM passes, both for a domain that is not ours. This is what a Google Workspace mailbox rule looks like from the outside when it forwards to an external address: the tenant rewrites the envelope sender to itself and signs with its own key, so both mechanisms pass for the tenant and neither aligns with the From domain. Ten different company domains appear this way in the period. They are not senders. They are the employers of people who subscribed to the newsletter and forward their mail somewhere else.
The classifier put both shapes under unknown, and on several of the tenant-forwarding rows the reason reads no_auth_pass, "nothing authenticated". Something did authenticate; it just did not align. The label is a one-word summary, and the row underneath it is the evidence, and here they point in different directions. That is survivable precisely because the row carries the authentication detail; a tool that returned only the label would have left the model with nothing to correct itself from.
"These ten companies are sending as you"
The model's draft named the ten tenant domains as sources sending mail as us. The rows above already say they are not: SPF and DKIM pass for each tenant's own domain, and neither aligns with ours. They should not be named at all. An aggregate report quietly reveals where your subscribers work, which is a good reason to keep a person between the model and anything published.
"Add the uncovered IPs to SPF"
get_spf_usage reports mail sent with the return path from IPs the record does not cover:
"uncovered": { "total": 888, "spfPass": 0, "dkimPass": 699, "dmarcAligned": 699 },
"uncoveredSources": [ { "ip": "209.85.220.69", "total": 565, "dkimPass": 518 }, ... ]
The model's draft suggested authorising them so SPF would pass. The largest is a Google relay, and 518 of its 565 messages have a passing aligned DKIM signature. These are forwarders, not senders. Adding them to the record would authorise Google's relay pool to send as the domain, which is worse than the problem it fixes, and DMARC is already passing on DKIM for most of that traffic. The dkimPass count sits inside every uncovered row for exactly this reason: it separates "a sender we forgot" from "a forwarder doing its job". The model had the number and did not use it until asked.
"Sending has stopped"
The anomaly report:
⚠️ Message volume 1 is below p5 quantile 10 (unusual drop)
"lastReportDate": "2026-09-08", "messagesQ95": 12889, "messagesQ5": 10

The domain page. Records health is Healthy, and activity health is Critical at the same time: a DNS check and a statistical check, kept apart on purpose. The two spikes are the entire sending pattern; no other day comes close.
The model's draft recommended checking whether the sending service had been disconnected. The timeline explains it: 9,440 messages on 26 August, 12,889 on 7 September, 463 on the day after the first send, and fewer than a hundred on every other day. This is a newsletter. It sends when there is an issue to send and is quiet the rest of the time, and on a distribution like that, the 5th percentile of daily volume is 10 messages, so the statistic is measuring noise. The flagged day is also the last day in the data, and aggregate reports for a day keep arriving for a day or two after it. That day has since filled in to 85 messages. lastReportDate sits in the same object as the flag.
The quantile is arithmetically correct. It is not an interpretation, and on a bursty sender it will not become one. The report can show that this domain sends in bursts; it cannot say why.
The screenshot at the top of this article comes from a second run that asks for the field behind each claim and the three-way sort, the first two items in the advice below, and says nothing about the send schedule. There the model reads the same two spikes as campaign-sized bursts rather than an outage. It got that from the shape of the timeline, not from any field, so it belongs in the inferred group, and it happens to be right. The prompt changed one thing: whether the reading was announced as a fact or filed as an inference. Confirming it still takes someone who knows the send schedule.
One line in that screenshot points at our own bug. Activity health runs once a day for the previous day and files the result under the date of the run, so its spikes land a day later than the timeline's, at whatever count had arrived by then. The model noticed the two series disagreed, checked which one reconciled with the summary, and went with the timeline. The labelling is ours to fix. The cross-check is the habit this article argues for.
The One That Stays Open
One row cannot be resolved from the data at all:
"isp": "Scaleway", "country": "NL", "total": 11,
"ipDomainName": "ultimategoalofcommunism.com",
"spfAuth": [{ "returnPath": "bounce.news.glockapps.co", "result": "softfail" }],
"dkimAuth": [{ "domain": "news.glockapps.co", "selector": "dkayeghju235...", "result": "permerror" }]

The one row the report cannot settle. The interface shows DKIM Fail; the API returns permerror, which means the signature could not be evaluated rather than checked and rejected.
Eleven messages, a hosting provider, a hostname chosen by someone with a sense of humour, and a DKIM permerror rather than a fail. Permerror means the verifier could not evaluate the signature at all, which is a different condition from a signature that was checked and did not match, and it does not distinguish a broken relay from a forged header. Eleven messages against 22,991 is not an incident. It is also the only row on this domain where "your own mail, mangled" and "someone else using your name" are both live options, and the report cannot choose between them. The record carries fo=1 and no ruf= address, so it asks for failure reports and gives them nowhere to go. Most large mailbox providers would not send them anyway. What would settle it is a copy of one of those messages with its headers, from whoever received it, and no DMARC report will ever contain that. This is the category where the model's job is to stop and say so.
How to Prompt It
Ask for the field behind each claim. "Which field is that percentage from?" is answerable when the tools return fields, and asking it exposes the inference.
Ask for findings sorted into read-from-DNS, read-from-reports, and inferred. The twelve percent figure and the spoofing reading both belong in the third group, and both announced themselves as first-group facts.
Ask what would confirm each inference. For the Scaleway row, the answer is one message with full headers from someone who received it. A model that says so is doing the job.
Say what a normal week looks like. A newsletter and a transactional sender have opposite baselines, and the reports cannot tell a model which one it is reading.
Do not ask it to write a DNS record from a report. A report says what failed. What the record should say encodes who is allowed to send, and that is a decision, not a lookup.
What This Is For
The model found nothing on this domain that I did not already know, which is the wrong test to apply. What it did was take two weeks of report data across 53 sources and, in a few minutes, arrive at the four questions a person has to answer: whether five months at p=none is still the intention, whether the mail that security gateways are breaking matters enough to change anything, whether the eleven Scaleway messages are worth chasing a sample for, and whether the ten forwarding tenants should be treated as a deliverability problem or ignored. Getting from raw XML to those four questions by hand takes an hour if you know what you are doing. The model does the hour. The four answers are still mine.
Postscript
Two days before the session this article is built on, DMARKOFF had sent me a notification about the same domain: authentication stable for more than fourteen days, alignment passing, current policy p=none, recommended policy p=quarantine. The product reached the same conclusion the model reached and the same one I reached, and it needed no LLM to get there. It is a threshold check, not an inference.

I have not acted on it. That is where this domain actually stands, and it is why this piece does not end on "the model analyses, and I decide". The decision was made months ago, when we chose the criteria that notification fires on. What was missing, I told myself, was five minutes and a DNS login.
So the question I end on is not the one I expected. Most of this article is about the ways a model should not be trusted to reason about DMARC data. The question is whether I would let something act on a decision I have already made and written down, while I am busy. Those are different kinds of trust. The first asks a model to judge. The second asks it to execute a rule I wrote.
I still have not clicked, and being busy is not the whole reason. Moving this domain to p=quarantine sends to spam the copies that security gateways mangle on the way into recipients' mailboxes, except where the receiver validates an ARC seal and trusts the original result. Google already did that for about 600 of the 1,389 failing messages in this period, reported as local_policy overrides with an arc=pass comment, so the open question is the remainder, and I have not decided whether I care about it. That is one of the four questions above, and it is genuinely open. When it closes, the change is mechanical, and I would probably let an agent make it.
The boundary that matters runs between decisions that have been made and decisions that have not, whoever carries them out. Everything here argues for keeping a person on the open ones. Almost none of it argues for keeping one on the closed ones.
Co-founder & CTO at DMARKOFF and GlockApps, email security specialist with 20+ years of experience, Golang & ClickHouse expert. A happy father of two teenagers.


