0%

← Back to course

Investigating a BEC Scam — with Claude Code

The case: Alice in finance received an urgent wire-transfer request from the CEO. The CEO’s office says he never sent it. Trace where the message actually came from and prove the sender is not who it claims to be — this time by putting the questions to Claude Code and checking what it does with them.

about 60 minutesClaude Code5 checkpoints
Steps
  1. •The brief
  2. 00 — Claude Code and evidence
  3. 11 — Parse the header
  4. 22 — Trace the delivery path
  5. 33 — Trace the originating IP
  6. 44 — Prove the forgery
  7. •Write the report

The brief

You are a junior investigator in the Cybercrime Unit of a corporate security team. The finance department has forwarded a suspicious email to the incident response portal. An employee, Alice, received a message that appears to be from the company’s CEO, asking for an urgent wire transfer to a new vendor in order to close a confidential deal. The amount is significant and the request is unusual.

The CEO’s office confirms he never sent it. The unit suspects Business Email Compromise — specifically the variety known as CEO fraud. Your job is a preliminary technical analysis of the message: trace the path it actually took across the internet, identify the infrastructure that sent it, and establish what makes it provably a forgery rather than merely a suspicious email.

The manual edition of this lab has you paste the header into a web parser and read it yourself. Here you keep the header on your own disk and put the questions to Claude Code instead: you ask, it picks the command, runs it, and shows you the output. The investigation is the same one. What changes is that you now have a second thing to examine — the tool’s own work.

The ground rule
Claude runs the commands. You own the findings. Never write down a value Claude did not show you the command for. A model will state a plausible domain name as readily as a true one, and it cannot be cross-examined — you can. Every prompt on this page therefore ends by asking for the command or the exact lines, and every step tells you how to re-run it yourself.
Where the risk is, this week
Email headers are the kind of evidence a model is most likely to get confidently wrong. It has read millions of them, so it can produce a fluent, plausible-sounding trace without looking at yours — and one of the things it most often gets backwards is which end of the Received: chain came first. Step 2 is built around that. Treat every claim as a claim about a file you can open.
How this page works
Each step ends with a checkpoint. The factual questions are the same ones the manual edition asks, because the evidence has not changed; the reasoning questions are about your handling of the tool. A step only counts toward your progress once every one of its questions is right. Nothing is sent anywhere; your progress is stored in this browser.

The header below is a constructed sample, not a captured message. The domains example.com and the addresses inside it are reserved for documentation; the originating IP is a real address in a real hosting network, which is what makes the lookup in Step 3 worth doing.

0

0 — Claude Code and evidence

Before you start

You need Claude Code installed and signed in. If it is not, work through the setup guide first — it takes about thirty minutes and this lab assumes it is done. Check with claude --version; a version number means you are ready.

Make a folder and start Claude there

Claude Code works inside the folder you launch it in, so give the evidence a folder of its own.

bash
mkdir ~/Lab2-BEC
cd ~/Lab2-BEC
Windows PowerShell: mkdir $env:USERPROFILE\Lab2-BEC, then cd $env:USERPROFILE\Lab2-BEC.
Build the evidence file

Copy the block below in full, from Delivered-To: down to the closing </html>, and save it in that folder as bec_email_header.txt. Save it with Unix (LF) line endings and a single trailing newline, or the hash will not match. Do this part by hand, in an editor — not by asking Claude to write the file for you. Evidence you did not place yourself is evidence you cannot vouch for.

Downloading bec_email_header.txt straight into that folder counts as placing it yourself, and is the safer route: no editor gets the chance to convert your line endings. Letting Claude write the file out does not count.
bec_email_header.txt
Delivered-To: alice.employee@example.com
Received: by 2002:a05:620a:10c8:0:0:0:0 with SMTP id z8csp2406822qkk;
        Fri, 13 Aug 2025 10:02:59 -0700 (PDT)
X-Google-Smtp-Source: ABCs/1234567890abcdefghijklmno=
X-Received: by 2002:a5d:43c4:0:b0:220:b922:a5a9 with SMTP id x4-20020a5d43c4000000b00220b922a5a9mr9988748wrp.89.1660399379017;
        Fri, 13 Aug 2025 10:02:59 -0700 (PDT)
ARC-Seal: i=1; a=rsa-sha256; t=1660399379; cv=none;
        d=google.com; s=arc-20160816;
        b=somehashvalue
ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816;
        h=to:subject:message-id:from:date:mime-version:dkim-signature;
        bh=anotherhashvalue;
        b=yetanotherhashvalue
ARC-Authentication-Results: i=1; mx.google.com;
       dkim=pass header.i=@ceo-offices.com header.s=google header.b=AbCdEfG;
       spf=fail (google.com: domain of support@v-pshosting.net does not designate 198.54.117.211 as permitted sender) smtp.mailfrom=support@v-pshosting.net
Return-Path: <support@v-pshosting.net>
Received: from mail.v-pshosting.net (mail.v-pshosting.net. [198.54.117.211])
        by mx.google.com with ESMTP id a6-20020a170902c30600b001a1d50c1157si910115plg.437.2025.08.13.10.02.58
        for <alice.employee@example.com>;
        Fri, 13 Aug 2025 10:02:58 -0700 (PDT)
Received-SPF: fail (google.com: domain of support@v-pshosting.net does not designate 198.54.117.211 as permitted sender) client-ip=198.54.117.211;
Authentication-Results: mx.google.com;
       dkim=pass header.i=@ceo-offices.com header.s=google header.b=AbCdEfG;
       spf=fail (google.com: domain of support@v-pshosting.net does not designate 198.54.117.211 as permitted sender) smtp.mailfrom=support@v-pshosting.net
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ceo-offices.com; s=google; h=mime-version:date:from:message-id:subject:to; bh=anotherhashvalue; b=somehashvalue
MIME-Version: 1.0
Date: Fri, 13 Aug 2025 17:02:55 +0000
From: "CEO John Doe" <john.doe@example.com>
Message-ID: <123456789.abscd@ceo-offices.com>
Subject: Urgent: Confidential Payment Request
To: alice.employee@example.com
Content-Type: text/html; charset=UTF-8

<html><body><p>Alice, I need you to process an urgent wire transfer for $15,000 to vendor account 87654321. This is for the confidential Project X. Please process immediately and confirm once sent. Discretion is key.</p><p>Regards,<br>John Doe<br>CEO</p></body></html>
In TextEdit on macOS, switch to plain text first (Format › Make Plain Text). In Notepad on Windows, save with Encoding: UTF-8 and not UTF-8 with BOM. A byte-order mark changes the hash.
Start Claude and verify the evidence

Now start Claude in that folder, and hash the file before anything else touches it. A hash proves the copy on your disk is byte-for-byte the copy that was published — a routine step in any forensic process, not a formality.

bash
claude
Prompt
Work out the SHA256 hash of the file bec_email_header.txt. Show me the command you ran and the output exactly as it printed. Do not change the file.

What Claude should do: run one hashing command and show you its output. Claude asks permission before it runs anything; read what it proposes before you press Enter. That habit matters more here than anywhere else — you are about to let a tool loose on evidence.

Check its work: run the command yourself and compare, character for character. This is the one value in the lab that certifies every other value, so it is the last one you should take on trust.

bash
# Linux, WSL
sha256sum bec_email_header.txt

# macOS
shasum -a 256 bec_email_header.txt

# Windows command prompt
certutil -hashfile bec_email_header.txt SHA256

Expected value:

sha256
ebabf124b0b1c4b747efd26f271f8b60101406be3b05b1bde8e22182504ae962
If your hash does not match, stop. Do not carry on and do not ask Claude to explain the difference away. A copy that does not hash to the expected value is not the evidence, and nothing you derive from it can be tied back to what was collected. Re-copy the block, look for a stray blank line at the end or line endings your editor converted to CRLF, and hash it again.

Checkpoint 0

Answer both to complete this step.

  1. What is the SHA256 hash of your copy of bec_email_header.txt?

  2. Claude tells you the hash matches the published value, but does not show the command it ran. What do you do before recording it?

1

1 — Parse the header

Start with the shape of the message as a whole: who it claims to be from, who it went to, when, and how many servers touched it. Ask one question at a time — a prompt that asks for five things at once gets you a summary, and a summary is where values go missing.

The summary fields
Prompt
Read bec_email_header.txt. Show me these header lines exactly as they appear in the file, one per line and nothing added: From, To, Subject, Date, Return-Path, Message-ID.

What Claude should do: print six lines copied out of the file. Ask for them verbatim, as above. If it paraphrases — “the message is from the CEO” — ask again for the raw lines. You want the text as it sits on disk, because in this lab the exact spelling of a domain is the whole finding.

Count the hops
Prompt
How many lines in bec_email_header.txt start with "Received:" at the very beginning of the line? Do not count Received-SPF or X-Received. Show me the command and the line numbers.

Check its work: the anchor is what makes this right. Run it yourself:

bash
grep -c "^Received:" bec_email_header.txt
grep -n "^Received:" bec_email_header.txt
A model asked “how many hops?” without an anchored pattern will sometimes count the indented continuation lines, or fold X-Received: into the total. Both give you three or four hops instead of two, and a wrong hop count changes the story from “injected straight into the gateway” to “relayed through a chain”.
The authentication verdict
Prompt
Show me every line in bec_email_header.txt that contains "spf=" or "dkim=", exactly as written. Do not summarise them.

What to look for: an SPF failure, spelled out with its reason. Note it and leave it there. Read the reason text closely before you decide what it proves — which address does it name? The checkpoint below turns on that question, and Step 4 is where it becomes the case.

Checkpoint 1

Both must be right to complete this step.

  1. How many Received: headers does the message carry?

  2. Claude reports that the message “fails SPF, so the sender is spoofed”. What has the SPF failure actually established?

2

2 — Trace the delivery path

Each Received: header is added at the top of the message by the server that has just accepted it. So the newest stamp sits highest and the chain reads bottom to top in chronological order: the bottom-most Received: line is the earliest hop the message records, the top-most is the last.

Reading order
Hop numbering in the standard parsers runs the other way round from the file: Hop 1 is the earliest — the message arriving at mx.google.com from outside — because the tool has already reversed the raw header into time order. Read the raw text bottom-up; read a hop table top-down. Getting this backwards inverts the whole trace, and it is the mistake to watch for in what Claude hands you.
Ask for the order explicitly
Prompt
Show me both Received: headers from bec_email_header.txt with their continuation lines and their line numbers. For each one, tell me which server wrote it and what time it recorded. Do not put them in any order yet.
Prompt
Which of those two Received: headers records the earlier event, and why? Answer using the timestamps in the file, not the order they are printed in.

Check its work: two questions, deliberately, and the second one asks for the reasoning rather than the answer. Then read the lines yourself:

bash
grep -A3 "^Received:" bec_email_header.txt
  • The lower one — Received: from mail.v-pshosting.net ... by mx.google.com — records the moment Google’s mail server accepted the message from an external host. This is the earliest hop, and the one your investigation rests on.
  • The upper one — Received: by 2002:a05:620a:10c8... — is internal to Google’s own infrastructure. It happened later and it tells you nothing whatsoever about the sender.
This is the step where a fluent wrong answer costs you most. If Claude calls the top line the origin, or describes the chain as reading top to bottom in time, it has produced a trace with the sender and the recipient swapped. Do not correct it and move on — work out from the timestamps which line came first, satisfy yourself, and only then continue. The order is the finding.
Find the boundary of trust

The hop that counts is the earliest one added by a server you trust. Here that is the line written by mx.google.com. Everything at or above it was recorded by infrastructure under the receiving provider’s control and can be relied on. Anything below it would have been written by servers the sender may control, and could be fabricated in full — forged Received: lines are a standard way of padding a fake message with a plausible history.

What to look for: in this message there is nothing below the trust boundary at all. That absence is itself informative. The sender did not relay through a chain of hosts; it opened a connection straight to Google’s gateway, which makes mail.v-pshosting.net both the first hop and the earliest infrastructure the message can attest to.

Prompt
The Date: header and the timestamp in the mx.google.com Received: line use different time zones. Convert both to UTC and tell me whether they agree. Show your working.

Check that conversion by hand. It is two lines of arithmetic, the report asks for both times in UTC, and time-zone arithmetic is a place where a confident wrong answer is easy to miss.

Checkpoint 2

All three must be right to complete this step.

  1. At the earliest hop, which server did mx.google.com receive the message from?

  2. Of the two Received: lines in the file, which one records the earliest event?

  3. Claude lists the hops for you in what it calls chronological order. What makes that ordering safe to put in a report?

3

3 — Trace the originating IP

The earliest trusted Received: header carries the IP address of the machine that delivered the message to the gateway. That address is your primary lead on the attacker’s infrastructure.

Extract the address
Prompt
In the earliest Received: header, there is an IP address in square brackets. Show me that line and tell me the address. Then show me every other line in the file where the same address appears.

What to look for: the address should turn up twice — once in the Received: line and once as client-ip= in the Received-SPF: line. Two independent records of the same value in the same message is a small corroboration, and it costs you one extra question to have it.

Why the brackets matter: the hostname before them is supplied by the connecting client and can say anything it likes. The bracketed address is what the receiving server actually saw the connection arrive from, and it cannot be forged over a completed TCP connection — the handshake would not finish.

This is the opposite of the situation in Lab 1, where the source addresses were forged wholesale because a SYN flood never needs a reply. Here the sender wanted the message delivered, so the address is genuine. Say so explicitly in your report — it is the difference between a lead and a dead end.
Find out who operates it
Prompt
Run a whois lookup on 198.54.117.211 and show me the raw output. Then point out the organisation that owns the network, the AS number, and the abuse contact email address.
Prompt
What is the reverse DNS name for 198.54.117.211? Show me the command and its output, and tell me whether it matches the hostname the sending server announced in the header.

Check its work: both of these hit the network, so both are things you can reproduce exactly.

bash
whois 198.54.117.211
dig -x 198.54.117.211 +short

What to look for: a commercial hosting provider’s network in the United States, allocated to Namecheap under AS22612, with reverse DNS pointing at the provider’s own infrastructure (parkingpage.namecheap.com) rather than at mail.v-pshosting.net. That mismatch between forward claim and reverse reality is worth a line in the report on its own.

The value you actually need from this lookup is the abuse contact. That is the party a preservation request would be served on, and the party who holds the account records, payment details and access logs for the host. Everything else in the record is context.

A model asked where an IP address is will answer with a city, because that is the shape of the expected answer. Be careful what you take from it. A datacentre city is the location of a rented server, not of a person. Whoever sent this could have been anywhere, reaching that host over SSH or a web panel, and the hosting account was very likely paid for with stolen or anonymous payment details. “The attack originated in the United States” is a sentence that will not survive being questioned — and it is a sentence a summary will hand you unprompted.

Checkpoint 3

All three must be right to complete this step.

  1. What is the originating IP address?

  2. Which company operates the network that address belongs to?

  3. Claude reports the city the lookup returned. What does that city actually tell you?

4

4 — Prove the forgery

Now correlate. A forgery is proved by contradiction: the identities the message asserts set against the identities that were actually verified. Three domains appear in this message and no two of them agree. Ask about them one at a time, so that each answer is a line you can point at.

Which domain did each check authenticate?
Prompt
In bec_email_header.txt, what is the SPF result, and which domain did SPF check? Quote the line it comes from.
Prompt
What is the DKIM result, and which domain signed the message? Quote the d= tag from the DKIM-Signature line.
Prompt
Make a small table with three rows: the domain in the From: header, the domain SPF checked, and the domain DKIM signed. Put the exact domain in each row. Do not draw a conclusion.

Check its work: one command produces every line those three prompts draw on.

bash
grep -n "spf=\|dkim=\|^Return-Path:\|^From:\|^DKIM-Signature:\|^Message-ID:" bec_email_header.txt
The third prompt ends with “do not draw a conclusion” on purpose. Laying the three domains side by side and seeing that none of them match is the reasoning this whole lab is built around. If the tool hands you the conclusion, you have the right answer and none of the understanding, and the report will show it.
What the results actually mean
  • SPF is fail: the operator of v-pshosting.net published an SPF record and 198.54.117.211 is not in it. The sending host was not authorised even for the throwaway domain it put on the envelope.
  • DKIM is pass, on d=ceo-offices.com. Read that slowly, because a “pass” is easy to misread as a verdict of legitimacy — and a model summarising the header will often print it as a green tick. All it establishes is that whoever signed the message controls DNS for ceo-offices.com and that the signed headers were not altered in transit. It says nothing at all about example.com. ceo-offices.com is a lookalike domain the attacker registered, and a valid signature from a domain you own is trivial to obtain.
  • The mismatch between From: and Return-Path: is real, but SMTP permits it and mailing lists depend on it, so it is not proof of forgery on its own. What makes it damning here is its company: an urgent, confidential, unverifiable wire transfer, requested by a CEO whose office denies sending it.
  • The Message-ID is <123456789.abscd@ceo-offices.com>. A message genuinely composed on example.com’s mail system would carry a Message-ID generated by that system, in that domain. This one names the same third domain that signed it, which places the message’s origin outside the company entirely.
The decisive finding
Alignment
The From: domain is example.com. SPF authenticated v-pshosting.net and failed. DKIM authenticated ceo-offices.com and passed. Neither authenticated example.com. Not one identity check in this message vouches for the domain shown to Alice, so the claim that the CEO sent it has no technical support whatsoever. That non-alignment — and not the SPF failure by itself — is what makes the message provably a forgery. It is also precisely what DMARC evaluates, and why a published DMARC policy on example.com would have stopped this message at the gateway.

Checkpoint 4

All five must be right to complete this step.

  1. What is the SPF result?

  2. Which domain did SPF actually authenticate?

  3. Which domain signed the message with DKIM?

  4. What makes this message provably a forgery?

  5. Claude states the alignment conclusion for you, and states it correctly. What do you still have to do before it goes in your report?

Write the report

The checkpoints above establish the facts. The graded deliverable is the report you write from them — a preliminary analysis of the kind you would hand your incident response manager. Use the template in the PDF handout, which has the full structure and the blanks to fill.

It has four parts:

  • Evidence details — filename, SHA256, and both message timestamps: the time in the Date: header and the time the gateway stamped in the earliest trusted Received: header. Give both in UTC and say whether they agree.
  • Executive summary — three or four sentences for a manager who will read nothing else: what was attempted, against whom, by what mechanism, and why identifying the person responsible is not straightforward. No header-level detail here.
  • Key findings — the attack type with the evidence that establishes it; the target, which is a person and a finance process rather than a host, so there is no victim IP to report; the attack source with its network owner and reverse DNS; the technical indicators (SPF result and domain, DKIM result and signing domain, the From: domain, the originating IP, the hop count); the alignment finding in one sentence; and the social-engineering markers — urgency, confidentiality, authority, a request to bypass normal verification.
  • Attribution challenges — what the source address can and cannot establish, that the sending host may itself be compromised, that the domains were registered for this purpose, and what the next step would be: whose logs, whose records, through what legal mechanism, and how fast it would have to be taken.

You may use Claude to check your own draft, which is a different thing from having it write one:

Prompt
Read my report. List every claim in it that is not backed by a line you can point to in bec_email_header.txt or by a lookup I ran. Do not rewrite anything.
Write it in your own words. Every value you need is one you produced yourself in the five checkpoints, and a finding stated without the evidence behind it is not a finding — whether it was you or the tool that stated it.