0%

← Back to course

Investigating a Business Email Compromise (BEC) Scam

The case: Alice in finance received an urgent wire-transfer request from the CEO. The CEO’s office says he never sent it. Read the message’s headers, trace where it actually came from, and prove that the sender is not who the message claims.

about 60 minutesa browser and a terminal5 checkpoints
Steps
  1. •The brief
  2. 00 — Tools and evidence
  3. 11 — Parse the header
  4. 22 — Trace the delivery path
  5. 33 — Trace the originating IP
  6. 44 — Prove the forgery
  7. •Write the report

The brief

You are a junior investigator in the Cybercrime Unit of a corporate security team. The finance department has forwarded a suspicious email to the incident response portal. An employee, Alice, received a message that appears to be from the company’s CEO, asking for an urgent wire transfer to a new vendor in order to close a confidential deal. The amount is significant and the request is unusual.

The CEO’s office confirms he never sent it. The unit suspects Business Email Compromise — specifically the variety known as CEO fraud. Your job is a preliminary technical analysis of the message: trace the path it actually took across the internet, identify the infrastructure that sent it, and establish what makes it provably a forgery rather than merely a suspicious email.

Every value you need is in the message’s headers, which are a kind of postal history: each server that handles a message stamps it on the way past. Nobody deletes those stamps, because deleting them would break delivery.

How this page works
Each step ends with a checkpoint. Answer it from the evidence — every question tells you where in the raw header, or in which tool, the answer appears. A step only counts toward your progress once every one of its checkpoint questions is right. Nothing is sent anywhere; your progress is stored in this browser.

The header below is a constructed sample, not a captured message. The domains example.com and the addresses inside it are reserved for documentation; the originating IP is a real address in a real hosting network, which is what makes the lookup in Step 3 worth doing.

0

0 — Tools and evidence

The tools

Nothing to install. This lab runs on two web parsers and a terminal, and the terminal is used only to hash the evidence file and, optionally, to run a WHOIS lookup.

  • Google Admin Toolbox Messageheader — parses a raw header and lays the delivery path out hop by hop, with the delay each hop introduced.
  • MXToolbox Email Header Analyzer — a second parser, worth running as a cross-check. It also resolves the sending domain’s mail-related DNS records.
  • ipinfo.io (or geoiptool.com) — reports the approximate location of an IP address and, far more usefully for an investigation, the network that owns it.
A parser tells you what the headers say. It does not tell you whether they can be believed. Every finding on this page gets checked against the raw text afterwards.
Build the evidence file

Copy the block below in full, from Delivered-To: down to the closing </html>, and save it as bec_email_header.txt in a folder of its own — for example ~/Lab2-BEC. Save it with Unix (LF) line endings and a single trailing newline, or the hash will not match.

Rather not copy by hand? Download bec_email_header.txt and save it into that folder instead. Hash it either way — the check below is what tells you your copy is the evidence.
bec_email_header.txt
Delivered-To: alice.employee@example.com
Received: by 2002:a05:620a:10c8:0:0:0:0 with SMTP id z8csp2406822qkk;
        Fri, 13 Aug 2025 10:02:59 -0700 (PDT)
X-Google-Smtp-Source: ABCs/1234567890abcdefghijklmno=
X-Received: by 2002:a5d:43c4:0:b0:220:b922:a5a9 with SMTP id x4-20020a5d43c4000000b00220b922a5a9mr9988748wrp.89.1660399379017;
        Fri, 13 Aug 2025 10:02:59 -0700 (PDT)
ARC-Seal: i=1; a=rsa-sha256; t=1660399379; cv=none;
        d=google.com; s=arc-20160816;
        b=somehashvalue
ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816;
        h=to:subject:message-id:from:date:mime-version:dkim-signature;
        bh=anotherhashvalue;
        b=yetanotherhashvalue
ARC-Authentication-Results: i=1; mx.google.com;
       dkim=pass header.i=@ceo-offices.com header.s=google header.b=AbCdEfG;
       spf=fail (google.com: domain of support@v-pshosting.net does not designate 198.54.117.211 as permitted sender) smtp.mailfrom=support@v-pshosting.net
Return-Path: <support@v-pshosting.net>
Received: from mail.v-pshosting.net (mail.v-pshosting.net. [198.54.117.211])
        by mx.google.com with ESMTP id a6-20020a170902c30600b001a1d50c1157si910115plg.437.2025.08.13.10.02.58
        for <alice.employee@example.com>;
        Fri, 13 Aug 2025 10:02:58 -0700 (PDT)
Received-SPF: fail (google.com: domain of support@v-pshosting.net does not designate 198.54.117.211 as permitted sender) client-ip=198.54.117.211;
Authentication-Results: mx.google.com;
       dkim=pass header.i=@ceo-offices.com header.s=google header.b=AbCdEfG;
       spf=fail (google.com: domain of support@v-pshosting.net does not designate 198.54.117.211 as permitted sender) smtp.mailfrom=support@v-pshosting.net
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ceo-offices.com; s=google; h=mime-version:date:from:message-id:subject:to; bh=anotherhashvalue; b=somehashvalue
MIME-Version: 1.0
Date: Fri, 13 Aug 2025 17:02:55 +0000
From: "CEO John Doe" <john.doe@example.com>
Message-ID: <123456789.abscd@ceo-offices.com>
Subject: Urgent: Confidential Payment Request
To: alice.employee@example.com
Content-Type: text/html; charset=UTF-8

<html><body><p>Alice, I need you to process an urgent wire transfer for $15,000 to vendor account 87654321. This is for the confidential Project X. Please process immediately and confirm once sent. Discretion is key.</p><p>Regards,<br>John Doe<br>CEO</p></body></html>
Notepad on Windows and TextEdit on macOS both need care. In TextEdit, switch to plain text first (Format › Make Plain Text). In Notepad, save with Encoding: UTF-8 and not UTF-8 with BOM. A byte-order mark changes the hash.
Verify the evidence

Hash the file before you analyse it. In a real investigation this is what proves that the working copy is the evidence and has not changed since it was collected — a routine step in any forensic process, not a formality.

bash
# Linux, WSL
sha256sum bec_email_header.txt

# macOS
shasum -a 256 bec_email_header.txt

# Windows command prompt
certutil -hashfile bec_email_header.txt SHA256

Expected value:

sha256
ebabf124b0b1c4b747efd26f271f8b60101406be3b05b1bde8e22182504ae962
If your hash does not match, stop. Do not carry on and do not “note the discrepancy” in the report. A copy that does not hash to the expected value is not the evidence, and nothing you derive from it can be tied back to what was collected. Re-copy the block, look for a stray blank line at the end or line endings your editor converted to CRLF, and hash it again.

Checkpoint 0

Answer both to complete this step.

  1. What is the SHA256 hash of your copy of bec_email_header.txt?

  2. Suppose your hash had not matched. What is the correct next action?

1

1 — Parse the header

Start by getting the raw text into a readable shape. A parser lays out the delivery path and flags authentication problems automatically, which saves you reading forty lines of continuation-folded header by eye.

Run it through Messageheader
  • Open bec_email_header.txt and copy its entire contents (Ctrl+A, Ctrl+C, or Cmd+A, Cmd+C).
  • Go to Google Admin Toolbox Messageheader, paste into the box, and press ANALYZE THE HEADER ABOVE.
  • Repeat with MXToolbox. Two parsers agreeing is worth more than one parser asserting.

What to look for: the tool prints a summary table — message ID, subject, creation date, sender, recipient — and then a hop-by-hop table with the delay at each hop. Two things should stand out at once. The chain is very short: count the Received: headers. And the summary reports an SPF failure, in red. Note both and leave them; they are the business of Steps 2 and 4.

Count the hops in the raw text

Before you trust the tool’s hop table, confirm the count yourself. Every Received: header starts at column one; the indented lines below each one are continuations of it, not separate hops.

bash
grep -c "^Received:" bec_email_header.txt
grep -n "^Received:" bec_email_header.txt
^Received: and not Received. The anchor is what excludes the continuation lines, and it also excludes Received-SPF: and X-Received:, neither of which is a hop.

Checkpoint 1

Both must be right to complete this step.

  1. How many Received: headers does the message carry?

  2. The tool flags the message with a red SPF Fail. What has that failure established, on its own?

2

2 — Trace the delivery path

Each Received: header is added at the top of the message by the server that has just accepted it. So the newest stamp sits highest and the chain reads bottom to top in chronological order: the bottom-most Received: line is the earliest hop the message records, the top-most is the last.

Reading order
The Messageheader tool has already reversed the raw header for you and presents the hops in time order, so Hop 1 is the earliest — the message arriving at mx.google.com from outside. The final row is the last internal handoff before delivery to Alice. Read the raw text bottom-up; read the tool’s table top-down. Getting this backwards inverts the whole trace.
Read the two hops
bash
# each Received: header with its continuation lines
grep -A3 "^Received:" bec_email_header.txt
  • The lower one — Received: from mail.v-pshosting.net ... by mx.google.com — records the moment Google’s mail server accepted the message from an external host. This is the earliest hop, and the one your investigation rests on.
  • The upper one — Received: by 2002:a05:620a:10c8... — is internal to Google’s own infrastructure. It happened later and it tells you nothing whatsoever about the sender.
Find the boundary of trust

The hop that counts is the earliest one added by a server you trust. Here that is the line written by mx.google.com. Everything at or above it was recorded by infrastructure under the receiving provider’s control and can be relied on. Anything below it would have been written by servers the sender may control, and could be fabricated in full — forged Received: lines are a standard way of padding a fake message with a plausible history.

What to look for: in this message there is nothing below the trust boundary at all. That absence is itself informative. The sender did not relay through a chain of hosts; it opened a connection straight to Google’s gateway, which makes mail.v-pshosting.net both the first hop and the earliest infrastructure the message can attest to.

Compare the two timestamps while you are here: the Date: header says Fri, 13 Aug 2025 17:02:55 +0000 and the gateway stamped the message at 10:02:58 -0700. Convert both to UTC before you decide whether they agree; the report asks for this.

Checkpoint 2

All three must be right to complete this step.

  1. At Hop 1, which server did mx.google.com receive the message from?

  2. Of the two Received: lines in the raw text, which one records the earliest event?

  3. Why is the hop written by mx.google.com the earliest one you can rely on?

3

3 — Trace the originating IP

The earliest trusted Received: header carries the IP address of the machine that delivered the message to the gateway. That address is your primary lead on the attacker’s infrastructure.

Extract the address

In the tool: open Hop 1 — the earliest hop. It labels the sending server’s name and its IP address separately.

In the raw text: find this line, and read what is inside the square brackets.

raw header
Received: from mail.v-pshosting.net (mail.v-pshosting.net. [198.54.117.211])

Why the brackets matter: the hostname before them is supplied by the connecting client and can say anything it likes. The bracketed address is what the receiving server actually saw the connection arrive from, and it cannot be forged over a completed TCP connection — the handshake would not finish. The Received-SPF: line records the same value independently as client-ip=198.54.117.211, which corroborates it.

This is the opposite of the situation in Lab 1, where the source addresses were forged wholesale because a SYN flood never needs a reply. Here the sender wanted the message delivered, so the address is genuine. Say so explicitly in your report — it is the difference between a lead and a dead end.
Find out who operates it

In a browser: paste 198.54.117.211 into ipinfo.io and run the lookup. Note the organisation, the autonomous system number, the country, and the reverse DNS name.

bash
# the registry record: owner, allocation, and the abuse contact
whois 198.54.117.211

# the reverse DNS name
dig -x 198.54.117.211 +short

What to look for: a commercial hosting provider’s network in the United States, allocated to Namecheap under AS22612, with reverse DNS pointing at the provider’s own infrastructure (parkingpage.namecheap.com) rather than at mail.v-pshosting.net. That mismatch between forward claim and reverse reality is worth a line in the report on its own.

The value you actually need from this lookup is the abuse contact. That is the party a preservation request would be served on, and the party who holds the account records, payment details and access logs for the host. Everything else in the record is context.

Be careful what you take from the city the tool reports. It is the location of a rented server, not of a person. Whoever sent this could have been anywhere, reaching that host over SSH or a web panel, and the hosting account was very likely paid for with stolen or anonymous payment details. “The attack originated in the United States” is a sentence that will not survive being questioned.

Checkpoint 3

All three must be right to complete this step.

  1. What is the originating IP address?

  2. Which company operates the network that address belongs to?

  3. The lookup reports a city. What does that city actually tell you?

4

4 — Prove the forgery

Now correlate. A forgery is proved by contradiction: the identities the message asserts set against the identities that were actually verified. Three domains appear in this message and no two of them agree.

The three identities
FieldValueWho wrote it
From:"CEO John Doe" <john.doe@example.com>The sender. Nothing verifies it by default, and it is the only sender identity Alice’s mail client displays.
Return-Path: / smtp.mailfrom<support@v-pshosting.net>The sender, in the SMTP envelope. This is the address SPF checks.
Message-ID:<123456789.abscd@ceo-offices.com>The mail system that composed the message.

On the mismatch: the displayed sender and the envelope sender sit in different domains. SMTP permits that, and mailing lists and bulk senders depend on it, so a From: / Return-Path: mismatch is not proof of forgery on its own. What makes it damning here is its company: an urgent, confidential, unverifiable wire transfer, requested by a CEO whose office denies sending it.

On the Message-ID: a message genuinely composed on example.com’s mail system would carry a Message-ID generated by that system, in that domain. This one names a third domain — and it is the same domain that signed the message with DKIM. That places the message’s origin outside the company entirely.

Read the authentication results
bash
grep -n "spf=\|dkim=\|^Return-Path:\|^From:\|^DKIM-Signature:" bec_email_header.txt
  • SPF is fail, with the reason spelled out: domain of support@v-pshosting.net does not designate 198.54.117.211 as permitted sender. The operator of v-pshosting.net published an SPF record and 198.54.117.211 is not in it. The sending host was not authorised even for the throwaway domain it put on the envelope.
  • DKIM is pass, with header.i=@ceo-offices.com and d=ceo-offices.com in the signature. Read that slowly, because a “pass” is easy to misread as a verdict of legitimacy. All it establishes is that whoever signed the message controls DNS for ceo-offices.com and that the signed headers were not altered in transit. It says nothing at all about example.com. ceo-offices.com is a lookalike domain the attacker registered, and a valid signature from a domain you own is trivial to obtain.
The decisive finding
Alignment
The From: domain is example.com. SPF authenticated v-pshosting.net and failed. DKIM authenticated ceo-offices.com and passed. Neither authenticated example.com. Not one identity check in this message vouches for the domain shown to Alice, so the claim that the CEO sent it has no technical support whatsoever. That non-alignment — and not the SPF failure by itself — is what makes the message provably a forgery. It is also precisely what DMARC evaluates, and why a published DMARC policy on example.com would have stopped this message at the gateway.

Checkpoint 4

All five must be right to complete this step.

  1. What is the SPF result?

  2. Which domain did SPF actually authenticate?

  3. Which domain signed the message with DKIM?

  4. What makes this message provably a forgery?

  5. Does the From: / Return-Path: mismatch prove spoofing by itself?

Write the report

The checkpoints above establish the facts. The graded deliverable is the report you write from them — a preliminary analysis of the kind you would hand your incident response manager. Use the template in the PDF handout, which has the full structure and the blanks to fill.

It has four parts:

  • Evidence details — filename, SHA256, and both message timestamps: the time in the Date: header and the time the gateway stamped in the earliest trusted Received: header. Give both in UTC and say whether they agree.
  • Executive summary — three or four sentences for a manager who will read nothing else: what was attempted, against whom, by what mechanism, and why identifying the person responsible is not straightforward. No header-level detail here.
  • Key findings — the attack type with the evidence that establishes it; the target, which is a person and a finance process rather than a host, so there is no victim IP to report; the attack source with its network owner and reverse DNS; the technical indicators (SPF result and domain, DKIM result and signing domain, the From: domain, the originating IP, the hop count); the alignment finding in one sentence; and the social-engineering markers — urgency, confidentiality, authority, a request to bypass normal verification.
  • Attribution challenges — what the source address can and cannot establish, that the sending host may itself be compromised, that the domains were registered for this purpose, and what the next step would be: whose logs, whose records, through what legal mechanism, and how fast it would have to be taken.
Write it in your own words. Every value you need is one you produced yourself in the five checkpoints, and a finding stated without the evidence behind it is not a finding.