# Tools That Fetch URLs: SSRF and Allow-Lists — LLM Application Security

Source: https://www.geekswithgeeks.com/en/llm-security/o-ssrf

> Stop a browsing tool from reaching internal systems.

## The server fetches on the attacker's behalf

A "fetch this URL" tool runs on **your server**, inside your network. If the model (or an injected instruction) can choose any URL, the attacker can make your server request internal services (`http://localhost`, a cloud metadata endpoint such as `169.254.169.254`, an admin panel) and read the answer: **server-side request forgery (SSRF)**. Defences: **allow-list** the hosts a tool may contact; require `https`; **parse the URL properly** and compare the real `hostname` (not a substring: `docs.example.com@evil.test` and `docs.example.com.evil.test` are traps); block **private and link-local IP ranges** after DNS resolution, and re-check redirects; limit ports, response size and time; and run the fetcher in an **isolated network segment** with no access to internal services or credentials.

## A URL allow-list that resists tricks, run

I ran this with plain Python 3 (standard library only). All attacks here are harmless demonstrations on local data, using no real systems. Only the plain https URL on an allowed host passes. Plain http, a look-alike subdomain, the `user@host` trick, the cloud metadata address and a non-standard port are all rejected. A real defence also checks resolved IP addresses and redirects.

```python
from urllib.parse import urlparse

ALLOWED_HOSTS = {"docs.example.com", "api.example.com"}

def allowed(url):
    u = urlparse(url)
    return u.scheme == "https" and u.hostname in ALLOWED_HOSTS and u.username is None and u.port in (None, 443)

for url in ["https://docs.example.com/page", "http://docs.example.com/page", "https://docs.example.com.evil.test/x",
            "https://docs.example.com@evil.test/x", "https://169.254.169.254/latest/meta-data", "https://docs.example.com:8443/x"]:
    print(f"{allowed(url)!s:5} {url}")

```

Output:

```
True  https://docs.example.com/page
False http://docs.example.com/page
False https://docs.example.com.evil.test/x
False https://docs.example.com@evil.test/x
False https://169.254.169.254/latest/meta-data
False https://docs.example.com:8443/x
```

## Re-check after redirects

An allowed host can redirect to an internal address. Validate every hop.

**Quiz:** Why compare the parsed hostname rather than check whether the URL "contains" an allowed domain?

- [ ] Substring checks are slower
- [x] Tricks like docs.example.com.evil.test contain the allowed name but go elsewhere
- [ ] Hostnames are not case sensitive only
- [ ] There is no difference

*Answer:* Tricks like docs.example.com.evil.test contain the allowed name but go elsewhere. Parse first, then compare the real host exactly.
