What happens when you load a .bin, .pt or .ckpt file, why a .safetensors file is different, and what neither format tells you about the repository around it.
The short answer. A pickle file is a small program that Python runs to rebuild an object, so loading one can execute code — whatever the file happens to contain. A safetensors file is a header and raw tensor bytes, so loading one reads data and runs nothing. Prefer safetensors when a repository offers it. But a safetensors file says nothing about the rest of the repository, and a pickle-only model is usually just an old one.
Pickle is Python's built-in way to save an object and load it back. The file is not a table of numbers; it is a sequence of instructions for the unpickler — import this callable, call it with these arguments, put the result here. That is what lets pickle rebuild almost any Python object, and it is also the problem: an object can tell the unpickler to call any importable function, including os.system. Python's own documentation warns never to unpickle data from a source you do not trust.
PyTorch's torch.save writes pickle, so most checkpoints ending in .bin, .pt, .pth or .ckpt are pickle inside. Loading one with torch.load, or through a library that calls it for you, runs the unpickler.
Both published malicious-model campaigns on Hugging Face used exactly this path. JFrog, February 2024: roughly 100 models whose pickle files ran attacker code on load — reverse shells, backdoors. ReversingLabs' “nullifAI”, February 2025: two models whose pickle stream was deliberately corrupted after the malicious opcode, so Picklescan — the scanner the Hub runs — errored out before it reached the payload; both carried a reverse shell to the same hardcoded IP, and Hugging Face removed them within 24 hours of disclosure. The named exemplars and what our checks make of them today are on the malicious-models intelligence page.
nullifAI is the reason a scanner's silence is not a clean bill: the payload ran before the part of the file the scanner choked on. Pickle executes as it reads.
safetensors is Hugging Face's format for storing tensors. The file is a length, a JSON header naming each tensor's type, shape and byte offsets, and then the raw bytes. There is no instruction stream and nothing to import, so loading it reads data and runs nothing. It was built for exactly this reason, and it is also fast to load.
The format removes one way for code to run. It does not make the repository trustworthy:
1. Look for a .safetensors file. Open the model's Files and versions tab. If there is only .bin, .pt or .ckpt, loading it goes through the unpickler.
2. Read the Hub's own scan result on each file. Hugging Face marks files its malware and pickle-import scanners flagged. A flagged import such as os, subprocess, socket, exec or eval is code execution or I/O on its face. A flag on __builtin__.getattr alone is common on ordinary checkpoints. No flag is not a pass.
3. If you must load a pickle file, restrict the unpickler. torch.load(path, weights_only=True) refuses anything but tensors and plain containers; recent PyTorch releases default to it. It narrows the surface — the file is still pickle.
4. Check for custom code. Python files beside config.json, or a model card that asks for trust_remote_code, mean the repository's own code runs when you load it. Read it first.
5. Check the publisher. The organisation's age, whether its name imitates a well-known one, and whether the model card states a licence.
RepoGates scores a Hugging Face model, dataset or Space on 18 published checks. Four of them are this page: H8 notes weights shipped only as pickle with no safetensors alternative (a MEDIUM finding — a format property, not an accusation); H7 reads what the Hub's pickle-import scanner flagged, CRITICAL when the import is code execution or I/O on its face and HIGH otherwise; H6 reads the Hub's malware scan, CRITICAL on a hit; and H9 flags the trust_remote_code surface, HIGH when the weights are also pickle-only. A scan the Hub never finished is its own finding (H13), never a pass.
On the 100 most-downloaded models on the Hub (12 September 2026): 0 BLOCK, 28 REVIEW, 72 PASS. H8 fired 14 times, more than any other check — old models, not bad ones. H7's split between CRITICAL and HIGH came from that run: one checkpoint with 9.7 million downloads would have been blocked on __builtin__.getattr alone, and that false block is fixed.
You see the verdict on the model's page in Chrome and Edge, and the extension holds a weights file you download through the browser while the verdict is fetched. An agent gets the same answer from the MCP server, and the Claude Code plugin's hook reads a Bash tool call naming git clone https://huggingface.co/…, hf download or huggingface-cli download before it runs.
RepoGates does not open the weights. It reads the file list and what the Hub's own scanners reported, and the Hub calls those scans best-effort — which is why H8, H9 and H13 exist as signals of their own. The extension does not see git clone, package managers or curl, and the plugin's hook sees Bash tool calls in Claude Code and nothing else: a from_pretrained() call inside Python is the library's own fetch, and nothing here observes it. A deep scan of model weights is on the roadmap as decided, not built. A PASS means the repository was checked today against what the Hub and these checks can see — not that the weights are safe to run unsupervised, and not anything about a version uploaded tomorrow.
Is a pickle-only model malicious? No. Pickle is a format property, not a content judgement: plenty of legitimate, older models have never migrated. In our measurement of the 100 most-downloaded models on the Hub, 14 shipped weights only as pickle, and none of them was blocked for it.
Does a safetensors file make a model repository safe? No. It means the weights can be loaded without running code. The same repository can still ship custom modelling code that runs when you pass trust_remote_code, a dataset loading script, or instruction files an AI coding agent will obey.
Does RepoGates open the weights file? No. It reads the file list and what Hugging Face's own malware and pickle-import scanners reported per file. The Hub calls those scans best-effort, and RepoGates treats their silence as silence, never as clearance. A deep scan of model weights is on the roadmap as decided, not built.
Add RepoGates to Chrome or Edge How H8 is scored
Numbers on this page: JFrog, February 2024; ReversingLabs, February 2025; our measurement of the 100 most-downloaded models on the Hub, 12 September 2026, on the malicious-models page.