
How to Learn Malware Analysis in 2026: From an Empty VM to Your First Sample Report
Most people who want to learn malware analysis open Ghidra in week one, see a screen full of mov and call, and quietly close it. A month later they are back to watching videos about it instead of doing it.
The problem is not difficulty. It is order. Malware analysis has a strict dependency chain, and Ghidra sits near the end of it, not the start.
This post is the chain, in order. Six steps, five to eight months of part time effort, and a checkpoint at each stage so you know whether to move on or stay put. It also covers the parts people skip: where to get samples safely, how to keep your host clean, and what the output of this work actually looks like when a SOC has to use it.
What malware analysis actually is
You take a suspicious file and answer three questions. What does it do, how do we detect it, and what has it already touched.
That output feeds three consumers: detection engineers who need signatures, incident responders who need a timeline, and threat intelligence teams who need to know which family and which campaign.
It is not penetration testing. Pentesting is about getting in. This is about taking something apart after it is already inside, usually while someone senior is asking when the report will be ready.
Who hires for this: SOC teams at tier 2 and above, DFIR consultancies, antivirus and EDR vendors, threat intelligence teams, and large banks with in house response capability.
What you need: 16 GB RAM if you can manage it, 8 GB at a stretch, and about 100 GB of free disk for virtual machines.
What makes it different from other security tracks: patience. One sample can take a full working day. That is normal, not slow.
Step 1: Build an isolated lab
Time needed: 1 to 2 days
This is the one step you must not improvise. A badly built lab does not fail quietly. It encrypts your actual files.
The machines
- Windows analysis VM. Windows 10 in VirtualBox or VMware. Defender off, automatic updates off, no antivirus. This is where samples run
- FLARE VM. A free installer from Mandiant that drops the entire Windows analysis toolkit onto that VM in one script
- REMnux. A Linux distribution built for malware analysis. Second machine, used for network simulation and static tooling
- Give each 4 GB RAM and 60 GB disk if your host allows it
Isolation rules
- Host only networking, or an internal network between the Windows VM and REMnux. Never bridged. Never NAT with live internet while detonating
- No shared folders, no shared clipboard, no drag and drop while a sample is live
- Turn off guest addition network features after setup. Some samples escape through them
- Move samples in as password protected archives. The researcher convention is the password
infected
Snapshots
- Clean snapshot after setup, and after every tool installation
- Restore to clean before every new sample, with no exceptions
- Name them properly, you will be switching constantly
Checkpoint: your Windows VM can reach REMnux, neither can reach your host or the internet, and you can restore a clean state in under a minute.
Common mistake: analysing on the host because setting up a VM felt slow. This is how people lose their photos, their dissertation and their client files in a single afternoon.
Step 2: Learn the fundamentals
Time needed: 4 to 6 weeks
You cannot reverse what you do not understand. This is the month that decides whether you finish this roadmap.
Windows internals
- Processes, threads, handles, and what process injection means at the API level
- DLLs, the loader, and why DLL search order is an attack path
- The registry, run keys, services, scheduled tasks, the standard persistence spots
- Win32 APIs worth recognising on sight:
CreateProcess, VirtualAlloc, WriteProcessMemory, CreateRemoteThread, RegSetValueEx, InternetOpenUrl
The PE file format
- DOS header, NT headers, section table, and what normally lives in each section
- Import and export tables, and how imports reveal capability before you run anything
- Entry point, relocations, resources, and the places packers hide payloads
Assembly and C
- x86 and x64 basics: registers, the stack, calling conventions, the instructions you will see constantly
- What loops, conditionals and function calls look like in disassembly
- Enough C to read decompiler output: pointers, structs, string handling
Checkpoint: given a PE file, you can name its sections, list the interesting imports, and say what those imports suggest, all before executing it.
Step 3: Static analysis
Time needed: 2 to 3 weeks
Static analysis is everything you can learn without running the file. It is free, it is safe, and it frequently answers most of the question. Squeeze it dry before detonating anything.
First pass
- Hashes: MD5, SHA1, SHA256. Check them against public intelligence sources before anything else
- Real file type, because extensions lie. Check magic bytes
- Strings, ASCII and Unicode. Hunt for URLs, IPs, paths, registry keys, mutex names, command lines
- Compile timestamp and language hints, keeping in mind both can be faked
PE inspection
- PEStudio for imports, resources and indicators in one view. Start here
- Detect It Easy for compiler and packer identification
- Entropy, because high entropy sections usually mean packing, encryption or an embedded payload
- Resources, where embedded executables, config blobs and decoy documents hide
Reading imports like a story
This is the skill that separates a checklist runner from an analyst.
- Networking APIs plus file writing plus a run key suggests a downloader
- Crypto APIs plus file enumeration plus shadow copy deletion suggests ransomware
- Keyboard hooks plus clipboard access plus HTTP POST suggests an infostealer
- Very few imports plus high entropy usually means packed, and the real imports only appear at runtime
Checkpoint: from static analysis alone you can write two sentences predicting the sample’s behaviour, then confirm or correct them in the next step. Getting this wrong is fine. Not attempting it is the mistake.
Step 4: Dynamic analysis
Time needed: 2 to 3 weeks
Now you run it and watch everything it touches. Behavioural evidence is what defenders can actually detect on, which makes this the step that produces most of your usable output.
The monitoring stack
- Procmon for file, registry, process and network events. Filter aggressively or you will drown in noise
- Process Hacker for the live process tree, injected memory regions, loaded modules and strings in memory
- Regshot to snapshot the registry before and after, then diff
- Autoruns to reveal persistence the sample created
- Wireshark with INetSim or FakeNet to capture traffic and fake the internet, so the sample gets replies without ever touching real infrastructure
What to record, every time
- Files created, modified and deleted, with full paths
- Registry keys written, especially persistence locations
- Processes spawned, plus any injection into legitimate binaries
- Network indicators: domains, IPs, URI paths, user agents, beacon intervals
- Mutexes, which make excellent detection signatures
When nothing happens
A sample that appears inert is usually telling you something.
- It may be checking for a VM, a debugger, a specific keyboard layout, or system uptime
- It may be sleeping. Some families wait minutes or hours before acting
- It may need a command line argument, or a specific export called directly
- Note the evasion itself. That behaviour is intelligence too
Checkpoint: you can produce a clean behavioural timeline listing files, registry keys, processes and network indicators, and then restore your snapshot without thinking about it.
Step 5: Reverse engineering
Time needed: ongoing, and it never really ends
Static and dynamic tell you what happened. Reversing tells you why, and answers what behaviour cannot: what is in the configuration, how the key is derived, what the second stage would have been.
The tools
- Ghidra. Free, includes a decompiler, and has the best free learning material. Start here in 2026
- IDA Free. Excellent disassembler and still the default in many teams. Worth learning later
- x64dbg. The debugger you will live in for unpacking and runtime inspection
- Scylla or PE-bear for dumping unpacked binaries and fixing imports
The workflow
- Find the real entry point, then follow the flow. Do not read top to bottom
- Rename functions and variables as you understand them. Your future self will thank you
- Break on interesting APIs instead of single stepping through thousands of instructions
- For packed samples, run until the payload decodes in memory, then dump and repair the import table
Anti analysis tricks to expect
- VM detection through registry artefacts, MAC prefixes, driver names and CPU feature checks
- Debugger detection:
IsDebuggerPresent, timing checks, exception based tricks
- String obfuscation and runtime API resolution, so a strings dump shows you nothing useful
- Long sleeps designed to outlast automated sandbox timeouts
Checkpoint: you have manually unpacked at least one packed sample and extracted its configuration, such as the command and control addresses.
This is the slow part of the field. One sample understood deeply teaches you more than fifty samples skimmed.
Step 6: Detection and reporting
Time needed: every sample, from now on
Analysis nobody can act on is a hobby. This step turns your findings into something a SOC deploys tomorrow morning, and it is the part that actually gets you hired.
Write detections
- YARA rules built from stable strings and byte patterns, not from things that change with every build
- Sigma rules for behavioural detection from logs: process creation patterns, registry writes, suspicious parent and child chains
- Network indicators with a note on how long each is likely to remain valid
- Test every rule against clean software too. A rule that fires on legitimate binaries is worse than no rule at all
Map to a framework
Map observed behaviour to MITRE ATT and CK techniques across defence evasion, persistence, credential access and exfiltration. This is how your analysis connects to controls the organisation already has.
Report structure that works
- Executive summary, three sentences a manager can read
- Sample identification: hashes, file type, size, first seen
- Capability summary, what it does, in order
- Indicators of compromise, grouped by type
- Detection opportunities, the rules you wrote
- Recommended actions, containment and remediation
Checkpoint: someone who has never seen the sample can hunt for it across an entire estate using only your document.
Where to get samples safely
You need real samples, from sources that expect you to handle them carefully.
- MalwareBazaar free, tagged by family, archives password protected with the standard researcher password
- theZoo and vx-underground curated archives, good for studying known families
- Practical Malware Analysis labs the book’s sample set, still the best structured training material in this field
- Public sandbox reports read analyses of samples you also examined, then compare conclusions with yours
Handling rules: download only inside the analysis VM, keep samples in password protected archives, never store them on the host, and never rename a sample to something harmless looking. Future you will double click it.
A realistic six month schedule
Based on eight to ten focused hours a week.
| Period | Focus |
| Weeks 1 to 2 | Lab build, isolation testing, tooling, first harmless test file |
| Weeks 3 to 8 | Windows internals, PE format, assembly basics, reading simple disassembly |
| Weeks 9 to 11 | Static analysis on ten real samples, written predictions for each |
| Weeks 12 to 14 | Dynamic analysis, full behavioural timelines, first IOC lists |
| Weeks 15 to 20 | Ghidra and x64dbg, manual unpacking, configuration extraction |
| Weeks 21 to 24 | YARA and Sigma rules, three complete reports, public writeups, start applying |
Where this leads
- SOC Analyst, tier 2. Escalated alerts, suspicious file triage, detection tuning. The most common entry point
- DFIR consultant. Incident response engagements, forensic timelines, client reporting under pressure
- Threat intelligence analyst. Family tracking, campaign attribution, reporting for defenders
- Detection engineer. Writing and maintaining the rules that catch these samples at scale
- Reverse engineer. Antivirus and EDR vendors, deep analysis of new families
- Threat researcher. Public research, blog posts, conference talks
Frequently asked questions
Do I need to know programming?
You need to read code more than write it. Enough C to follow decompiler output, enough Python to automate parsing and extraction, and a working grasp of x86 assembly. You are not becoming a software developer.
Is it safe to analyse malware on my laptop?
Only inside a properly isolated VM with host only networking, no shared folders and clean snapshots. On the host directly, never. Some families specifically look for shared folders and mapped drives.
Ghidra or IDA for a beginner?
Ghidra. Free, has a decompiler, excellent learning material. Pick up IDA later because many teams standardise on it.
How long until I am employable?
Five to eight months of consistent part time study to be credible for a tier 2 SOC or junior DFIR role, assuming you finish with real reports and detection rules you can show.
Is malware analysis legal?
Analysing samples in your own isolated lab is legal in most jurisdictions, including India. Deploying malware against systems you do not own, or distributing samples carelessly, is not.
The part that actually decides it
The roadmap is not the hard part. You found it in ten minutes and you will find four more like it today.
The hard part is week five, when you are reading about the PE header for the third time and none of it feels like malware analysis yet. The people who make it through are not smarter. They just kept going while everyone else went looking for a faster roadmap.
Build the lab this weekend. Pick one sample from MalwareBazaar next week. Write down what you think it does before you run it, then find out how wrong you were. Do that ten times and you will already be doing the job.
Publishing your sample analyses on Hacklido is a good way to keep yourself honest. Writing for readers forces the clarity that reports demand, and your writeups become the portfolio you will point interviewers to.