The SSRF advisory lands the way they always do: a feature that fetches a URL — webhook delivery, PDF preview, an avatar import — turns out to accept an internal address. The audit trail is short. An input field took http://169.254.169.254/latest/meta-data/, the instance metadata service answered, and the response made it back to the attacker inside a JSON error message. No exploit chain, no zero-day. Just an HTTP client that did exactly what it was told.
Server-side request forgery is one of those vulnerabilities that survives layered defenses because it abuses a trust boundary most architectures never draw: everything inside the VPC trusts everything else. Your app can talk to the database, the metadata service, the admin panel, and the payment processor — and by default, so can anything that controls where your app points its HTTP client. This post covers what actually defends against it in 2026: input validation that doesn’t lie to you, egress controls, the split DNS architecture big providers converged on, and the metadata service hardening that came out of the Capital One breach.
The four failure shapes
Nearly every SSRF incident falls into one of four patterns, and each demands a different defense:
- Basic SSRF — the response body is returned to the attacker. This is the loud, obvious case, and the reason “just don’t echo the response” is necessary but nowhere near sufficient.
- Blind SSRF — the response never reaches the attacker, but the request still happened. Side channels do the exfiltration: response timing, connection errors surfaced in logs the attacker can read, or an out-of-band DNS lookup to a domain they control. If you can detect that the server made a request, you have a channel.
- Partial-blind via errors — connection errors leak structure: “connection refused” versus “timeout” distinguishes an open port from a closed one, which is a port scanner running through your fetch endpoint.
- Chained SSRF — the internal service reached is itself vulnerable, and the attacker pivots: hit an unauthenticated admin endpoint, an internal API with no rate limiting, or the metadata service for credentials. The Capital One breach in 2019 was exactly this: SSRF to the instance metadata endpoint harvested IAM role credentials, which were then used against S3. Roughly 100 million records. The SSRF was the entry; the over-privileged metadata service was the payload.
Input validation: the defenses that look like they work
The naive fix — a denylist or a regex on the URL string — fails for a reason worth understanding, because the same confusion between strings and parsed authority breaks other URL security checks too.
Consider what a URL actually is. The WHATWG URL specification defines a parser whose job is to normalize a string into components: scheme, host, port, path. But parsing is not canonicalization, and browsers, standard libraries, and fetch libraries all resolve URLs slightly differently. That gap is where the bypasses live:
- Redirects — you validate
https://attacker.com/start, it responds302 → http://169.254.169.254/, and your client follows it. Validation happens once; the client navigates many times. - DNS rebinding — you resolve
attacker-controlled.example, get203.0.113.5, pass the check, then the client performs a second resolution at request time — and the authoritative DNS server has switched the answer to169.254.169.254. The check and the fetch use different DNS answers. - Alternative IP encodings —
2130706433,0x7f000001,0177.0.0.1, and127.0.0.1all parse to the same address depending on the parser. A validator that works on strings and a client that works on parsed addresses disagree. - URL parser confusion — libraries disagree about backslash handling in schemes, about what counts as a delimiter, about IPv6 zone IDs. A URL that validates to one host in your checker resolves to another in the client.
The OWASP SSRF Prevention Cheat Sheet is blunt about the implication: validation must operate on the resolved IP address, not the hostname string, and it must happen on every connection. The defensible implementation pattern is a two-step check: parse the URL, extract the host, resolve it yourself, then validate the resulting IPs against your policy — and then connect to the IP you validated, with the Host header set to the original hostname (so TLS and vhost routing still work). Resolving twice — once for validation, once inside the client — is what creates the rebinding window.
// DialContext that enforces policy on the resolved IP itself,
// closing the DNS-rebinding window (validate-then-connect to the
// same address).
func SafeDialContext(allowed func(net.IP) bool) func(ctx context.Context, network, addr string) (net.Conn, error) {
return func(ctx context.Context, network, addr string) (net.Conn, error) {
host, port, err := net.SplitHostPort(addr)
if err != nil {
return nil, err
}
ips, err := net.DefaultResolver.LookupIPAddr(ctx, host)
if err != nil {
return nil, err
}
var conn net.Conn
for _, ip := range ips {
if !allowed(ip.IP) {
continue
}
conn, err = new(net.Dialer).DialContext(ctx, network, net.JoinHostPort(ip.IP.String(), port))
if err == nil {
return conn, nil
}
}
if err != nil {
return nil, err
}
return nil, fmt.Errorf("no resolved address for %s passed egress policy", host)
}
}
func privateBlocked(ip net.IP) bool {
return !(ip.IsLoopback() || ip.IsPrivate() || ip.IsLinkLocalUnicast() || ip.IsUnspecified())
}
client := &http.Client{
Transport: &http.Transport{DialContext: SafeDialContext(privateBlocked)},
CheckRedirect: func(req *http.Request, via []*http.Request) error {
if len(via) >= 3 {
return errors.New("too many redirects")
}
return nil // SafeDialContext re-validates each hop's host
},
Timeout: 10 * time.Second,
}
The property that makes this work: the allowlist check runs on the IP inside the dialer, which is the same code path the client uses to open the socket. There’s no separate “validate” step that can drift out of sync with the “connect” step. Every redirect hop re-enters the dialer, so the policy holds across the whole navigation chain. Note that IP allow/denylists on parsed IPs don’t save you alone — the checks must apply per-hop, and the dialer is the only place that sees every hop.
Also worth fixing globally: block non-HTTP schemes (file://, gopher://, dict://) by checking the scheme after parse — and treat any http:// URL where the host is a bare IP in a private range as invalid, since a hostname is what an external target should have.
Egress controls: draw the boundary in the network
Application-level validation will eventually miss something — a new scheme, a parser bug, a code path someone forgot. Network egress policy is the layer that makes the miss survivable. The principle: workloads that have no business calling internal services shouldn’t be able to, regardless of what their code does.
- Separate the fetch workload. Run URL fetching in its own service, on its own nodes or subnet, with its own security group or NetworkPolicy. The main application tier never makes outbound requests to user-controlled URLs; only the fetcher does. When the fetcher is compromised, the blast radius is “the fetcher can reach the internet,” not “the fetcher can reach the database.”
- Default-deny egress for that tier, with an allowlist of destinations the feature actually needs (document formats, image CDNs, known partner webhooks). Kubernetes
NetworkPolicyegress rules or a service mesh can express this; a NAT gateway with a domain allowlist is a blunter but workable approximation. - Strip what the request doesn’t need. The fetcher should send no Authorization headers, no cookies, no session material of its own. If the request carries ambient credentials, SSRF becomes credential forwarding.
- Sanitize errors. Map all dial/DNS/TLS failures to a single generic message. “Connection refused to 10.0.4.12:5432” in an API response is reconnaissance handed over for free.
The fetch-service pattern also gives you a place to enforce response policy: cap response size before buffering, enforce content-type allowlists, and disable decompression bombs by setting limits on the decompressed stream, not the compressed bytes.
The IMDS lesson: assume SSRF, contain the blast radius
The single highest-leverage SSRF defense in cloud environments is instance metadata hardening, because it converts the worst outcome (credential theft) into a nuisance. Both major clouds moved after real incidents:
- AWS IMDSv2 — documented in the EC2 instance metadata configuration guide, IMDSv2 is session-based: the client PUTs to
/latest/api/tokenwith a hop limit of 1, receives a token, and presents it on subsequent requests. Plain GET — which is what every SSRF payload ships as — gets no data. The hop limit ensures the request can’t survive a proxy hop. Enforce it withHttpTokens=requiredso IMDSv1 is rejected outright;HttpPutResponseHopLimit=1blocks container escapes that traverse the host network stack. - GCP metadata headers — the metadata server requires
Metadata-Flavor: Googleon every request, and rejects requests that carry anX-Forwarded-Forheader (a standard artifact of having passed through a proxy). Both defenses specifically target the properties of an SSRF request rather than its destination. - Azure — requires the
Metadata: trueheader, same idea.
The deeper lesson generalizes: the metadata endpoint was designed assuming the caller was trusted because it was “inside the machine.” IMDSv2’s designers inverted the assumption — a request is untrusted unless it proves it originated on this host (TTL=1) with a client capable of a two-step protocol. Every internal endpoint that returns credentials deserves the same scrutiny: who can actually reach this, and what proves they’re allowed to?
Split DNS and the 2026 state of the art
The remaining hard problem is validation timing: the DNS rebinding window exists because validation and connection resolve separately. The architectures that close it fully converge on the same shape — the thing that validates is the thing that connects, over a single resolution:
- Pin the resolved IP into the connection. Resolve once, validate the IP, connect to that IP with SNI/Host set from the original URL (the Go dialer pattern above does this directly). No second resolution, no rebinding window.
- Split-horizon DNS on the fetch tier. Give the fetcher a DNS view where internal names don’t resolve at all. Then even a successful validation bypass — a parser bug, a missed scheme — fails at resolution, because the internal zone simply isn’t visible from that network segment.
- Proxy everything through an egress gateway. The fetcher has no direct network route anywhere; its HTTP client points at a forward proxy that enforces the destination policy. One component owns the policy, gets patched, gets logged. This is where the SaaS offerings in this space sit, and why — they centralize the single hardest thing to get right in every application separately.
One more 2026-relevant development: browser-side defenses have hardened around the same threat. The Origin header and Private Network Access (formerly CORS-RFC1918) now let servers distinguish requests coming from public websites versus private networks, and Chrome progressively restricts public sites from addressing private addresses. That closes the client-side SSRF variant — a victim’s browser being used to attack their router or intranet — but it does nothing for server-side SSRF. The two problems rhyme; the defenses are different layers.
A practical checklist
- Parse URLs; never match them with regex. Extract scheme, host, and resolved IPs.
- Validate scheme against an allowlist (usually just
https:, maybehttp:). - Resolve the host yourself and check the IP against deny ranges: loopback, RFC 1918, link-local
169.254.0.0/16, unique-localfc00::/7, link-localfe80::/10, and your cloud’s metadata address specifically. - Connect to the validated IP with a custom dialer — one resolution, validated and used.
- Cap redirects and re-validate per hop (the dialer gives you this for free).
- Run user-driven fetches in a dedicated tier with default-deny egress and no ambient credentials.
- Set
HttpTokens=requiredand hop limit 1 on every EC2 instance; verify the equivalents on GCP and Azure. - Return one generic error for every outbound failure. Log the details internally with the correlation ID, not in the response.
- Set response size caps and timeout floors on the fetch client before the first request goes out.
SSRF rewards defenses that assume your validation code will eventually be wrong. Validate carefully, but put the real wall in the network — a fetcher that can’t route to private space is a contained incident, while one that can is a breach waiting for whoever finds the endpoint first. The Capital One breach wasn’t caused by sophisticated exploitation; it was caused by an unfetchable URL being fetchable. Draw the boundary in the network, enforce the metadata service’s two-step handshake, and the worst an SSRF bug can do is fail loudly instead of silently handing over the keys.