Lesson 21 / 28
Models, Packages and Plugins as Dependencies
Apply normal supply-chain hygiene to model files and AI tooling.
You run other people's code and data
An LLM application depends on many third-party pieces: model files downloaded from hubs, Python and JavaScript packages, datasets, agent plugins and MCP servers, prompt templates and container images. Each can be malicious or compromised. Practical rules: download from trusted, verifiable sources; pin versions and verify hashes or signatures; scan dependencies; prefer safe serialisation formats (such as safetensors or plain JSON) over formats that can execute code when loaded; run unfamiliar models and tools in a sandbox; review third-party plugins and tools for the permissions they request; beware of look-alike names (typosquatting, including package names an AI assistant hallucinates, which attackers can register); and keep a bill of materials of the models, datasets and packages you use, with licences. Treat model providers as vendors: review their security and data terms.
Trust what you load; bound what you serve
Models, packages and plugins are dependencies; requests and tokens are resources an attacker can burn.
Loading an untrusted pickle runs code, run
I ran this with plain Python 3 (standard library only). All attacks here are harmless demonstrations on local data, using no real systems. Unpickling an object whose __reduce__ returns print executes that call during loading; here it is a harmless print, but a real attacker would run any command. Loading JSON, a data-only format, cannot run code. This is why pickle-based model files from untrusted sources are dangerous.
import pickle, json
class Payload: # a harmless stand-in for a malicious object
def __reduce__(self):
return (print, ("[!] code ran while LOADING the file",))
blob = pickle.dumps(Payload())
print("loading an untrusted pickle...")
pickle.loads(blob) # merely loading executes code
data = json.loads('{"weights": [0.1, 0.2]}') # data-only formats cannot run code on load
print("json loaded safely:", data)
Output:
loading an untrusted pickle...
[!] code ran while LOADING the file
json loaded safely: {'weights': [0.1, 0.2]}Keep a bill of materials
List models, datasets and packages with versions and licences so you can respond fast when one is compromised.
Quick check: Why is loading a pickle file from an unknown source risky?
- Pickle files are always too large
- Unpickling can execute arbitrary code during loading
- Pickle cannot store numbers
- It only slows down training
Answer
Unpickling can execute arbitrary code during loading — Prefer data-only formats and verify the source of any model file.