Guide
Leak Collector
Read a leak site’s file listing over Tor or the clearnet, tick what you need, take it away as one packed archive.
The short version
- Paste the address of the file listing — not the site’s front page.
- Wait for the crawl. It reads index pages and downloads nothing.
- Read the list, tick what you actually need, and tick Forensics beside anything you want analysed.
- Press Retrieve. Up to five files come down at once, with a progress bar each.
- Download the pack. One zip per run, with a manifest naming the source and SHA-256 of every file.
The idea
Ransomware groups publish their victims' documents behind a file listing — a tree view, an index, a grid — where you click one file at a time and repeat. That is the whole complaint this tool answers.
You point it at the listing. It reads the index, shows you everything on offer, you tick what you need, and it comes back as one archive with a manifest saying where every file came from.
What it does to this server, and why the gate is high
Everything you tick is fetched from this server's address and written to this server's disk, where it sits under this operator's name until you take the archive away. These are somebody else's documents, obtained from a criminal publication, and they remain somebody else's personal data the entire time this platform holds them. Every step is written to the audit log with your name on it. Have your legal basis before you tick, not after.
That is why the tool sits at Security Specialist and is not inherited from below. Every other tool here reads a record somebody published, or scans a host whose owner proved they consented. There is no equivalent proof available for this one — nobody can consent on behalf of the people whose documents are on that index — so the control is the account holding it.
The pause in the middle is the design
The lifecycle is:
scraping → listed → [you tick boxes] → queued → downloading → packing → ready
Nothing is fetched between listed and queued. The crawl reads index pages and writes down what was on offer; only you, ticking named boxes, causes a single document to be pulled.
That pause is the entire difference between retrieving eleven documents you named and mirroring a
ransomware group's whole victim estate from this address. It is also how this tool sits beside the
onion_probe source, which reads one page of a leak site and refuses to walk it: that
one runs unattended inside an automated fan-out against any onion an investigation happens to meet,
and nothing about that act has a person in it.
Pointing it at a listing
Paste the address of the file list, not the site's front page. Onion and clearnet both work and the tool picks the transport from the address.
- It never signs in to anything. A listing behind a login is a listing this tool cannot read, and that is the correct outcome. A URL with credentials in it is refused rather than silently stripped.
- It never leaves the host you pointed it at, and it never goes above the folder either. A link back up the tree, or sideways out of your folder into another victim's, is not followed at any depth — so one job cannot turn into an open-ended crawl from this address.
- A page with no links in it might still have a list. Some sites publish a viewer rather than an index and load the file list separately; the tool reads the address out of the page and fetches that instead. Nothing on the page is executed.
- Onion listings are fetched over port 80. Anything else is refused with the reason rather than left to time out looking like a dead site.
Depth, label and parallel downloads
- Depth — how far to follow sub-listings. 1 is the page you pasted, and there is an Unlimited option. On a real dump the documents are usually four or five levels down — Department / Person / Year / Case — so a shallow crawl reliably lists the folder above everything you came for. Reading index pages only; nothing is downloaded at any depth. Bounded by a page ceiling, a file ceiling, and the rule above that it never climbs out of your folder. Unlimited over Tor can run for an hour.
- Label — optional and worth writing. "Rhysida — client X, IR-2026-114" is findable in a year; a UUID is not.
- Parallel downloads — 1 to 5, used later. Five is about a hidden service's whole comfortable budget; a sixth mostly buys timeouts.
Reading the list — including when it is empty
The crawl summary is the first card on the job page and it is worth reading before the table, especially when the table is short. A crawl that found nothing tells you which of four things happened:
| No links at all | Almost always a JavaScript application: the file list arrives in a separate request. Open your browser's network tab, find that request, and paste its address instead. |
| Links, none kept | They pointed off-site, back up the tree, or at the page's own furniture. This is what an index that only links elsewhere looks like. |
| Sub-listings, no files | Raise the depth, or start a new job pointed at one of them. |
| Would not parse | The page did not load or is not HTML the parser can read. |
Only one of those means the documents are gone. "No files found" would have read as all four.
Sizes are claims until they are not
A size shown as ~2.4 MB is what the listing said. A size shown as 2.4 MB
is what actually arrived. They are kept side by side because a stated size that turns out wrong is
itself worth seeing. The ticked-total in the toolbar says "by the listing's own figures" for the
same reason.
Ticking and retrieving
- Get — retrieve this file.
- Forensics — also send it to the forensic tool. Ticking this ticks Get too, because analysing a file you are not retrieving is not a thing that can happen.
- Search, then Tick all matching — which means everything the filter matched across the whole collection, not just the page in front of you. It asks first and tells you the number.
- Tick this page when you mean the page. Ticks survive paging: tick forty on one page, search for something else, tick six more, and all forty-six go.
Finding the files, and finding the folder
The list is paged and the search runs over the whole collection, not the page you are
looking at. Every word has to appear somewhere in the full path or the name, in any order —
invoices 2009 pdf — and a leading minus excludes: garden -thumb. Type,
size and state narrow what the words left. The address bar carries all of it, so a filtered view is
a link you can paste into a ticket.
Often what you are hunting is a place rather than a file. Searching offers the folders it matched above the results — click one to see that folder and everything under it. A folder with no files in it is still offered, and says which of two things happened: the crawl never opened it, or it opened it and there was nothing there. Those need different actions, so they are not allowed to look the same.
Press Retrieve ticked files. Up to five come down at once with a progress bar each. A bar that sweeps rather than filling means the server declared no length — the byte count is in the tooltip.
Files that already downloaded are not fetched again; the confirmation says how many you ticked that are already here. Retrying a failed transfer is just ticking it again — a Tor circuit dying halfway is routine.
"check it" is not "done"
A row marked check it downloaded fine, but the server returned an HTML page for a file that is not HTML. That is very often an error page served with a 200 rather than the document. Nothing is refused — it might genuinely be HTML — but a pack full of identical 3 KB "not found" pages with every row saying done is the most plausible way this tool could mislead you, so the row says so.
Ceilings
There is a per-file limit and a per-job limit, both enforced on the bytes as they arrive rather than on what the server declared. Files past a ceiling are marked skipped with the reason — a job that ran out of budget did not find fewer files.
The pack
When the transfers settle, the set is zipped into archived/<date>-<site>.zip
and the loose files are deleted. Download it from the Packs card.
Individual files are deliberately not downloadable. Partly because nobody working an incident wants to click 340 times — but mostly because a per-file download route would make this platform a mirror of a ransomware group's estate with a better index and a login. One zip is one act, with one row against it, one count of how many times it left, and one name attached.
MANIFEST.txt
Every pack opens with one. It states the source address, the network, the job reference, who requested it, when it was packed, and for every file: its source URL, its byte count, its SHA-256 as written to disk, and the second it was fetched.
A folder of loose files with no provenance is worth very little to an investigation and is actively dangerous as evidence. The manifest is what lets the pack be handed on and still say what it is — in plain text, because it has to be readable in a year by somebody with a zip utility and quite possibly a lawyer rather than an engineer.
Running it again
You can go back to the list and tick more. Each run produces its own pack containing that run's files — the previous run's loose copies are already gone, which is the point of pruning them.
The forensic hand-off
Ticking Forensics hands the file to the forensic tool once it lands, and a link to the report appears on the row.
- The tick is the consent. The forensic tool's rule is that analysis never happens as a side effect of a file arriving. Ticking a named file before a single byte is fetched is a more deliberate act than pressing Analyse on an upload, not a less deliberate one.
- It runs after the pack, not during. Your zip is downloadable the moment it exists; the analyses land behind it. Interleaving them would stall a transfer slot for each one, so asking for analysis would delay a zip that was already on disk.
- A file too large is reported as skipped with its size, never silently unanalysed. The forensic side reads the whole file into memory and has its own ceiling.
Stopping, and what happens to the files
Stop cuts transfers in flight within about a second. Anything already downloaded is kept and can still be packed.
| Loose files | Deleted the moment the pack is written, and in any case within a day. They are scaffolding. |
| Packs | Kept until you or an administrator deletes them. No automatic expiry — an engagement outlives any number that would be safe to guess, and there is no backup. |
| The rows | Survive a prune. "This file was retrieved, it was 4.2 MB, here is its SHA-256" is the record of what happened. |
| Delete everything | Removes the job, every file and every pack. No undo, no backup. |
Delete your packs when you are finished with them. The volume is shared with the timeline and forensic tools, and nothing else on it writes gigabytes at a time. There is no expiry precisely so that the decision is a person's — which means it has to actually be made.
Finding your way around
Four sections down the left:
- Collect — point it at a listing and start.
- Queue — what is crawling, downloading or packing right now, with a progress bar per file and a Stop on each job. This is the page to leave open, or to come back to. A collection keeps going when you navigate away or close the tab, and the number beside Queue in the sidebar tells you from anywhere in the tool that something of yours is still running. It also lists what finished in the last few hours, so coming back to an empty queue tells you whether your pack was built or your job died.
- My collections — everything you have run, with a share toggle on each row.
- Org collections — what your team has shared.
Who else can see it, and sharing
Private by default. A collection is yours and platform administrators', and nobody else's, until you say otherwise. There is deliberately no public option: a collection is an index of somebody else's stolen documents, and there is no version of publishing that which is defensible.
Share with org — from a row in My collections, or from the collection itself — puts it under Org collections for everyone in your organisation.
Sharing does more than let colleagues look. Anyone in your organisation can then retrieve more files from that leak site and download its packs — spending this server's bandwidth against a collection you started. Every one of those actions is audited against whoever did it, not against you. That is the trade, and it is why the confirmation says so.
Two things stay with you: un-sharing and deleting. A colleague adding to the work is one thing; a colleague removing your evidence with no undo, or cutting off everyone else's access, is another.
A tick is a request, not a label
Both checkboxes clear once a file settles. They mean "do this on the next run", not "this file is a forensics file" — what actually happened is the state pill and the forensics line under the name, which stay put with a link to the report.
Worth knowing because it used to work the other way and was wrong: the Forensics tick persisted after a completed run, so the next batch you retrieved would quietly re-analyse the old file as well unless you remembered to clear it.