A colleague pinged me last week with a diff and one line of commentary: "why are there TWO of these now?" The diff swapped a parse_url() call in our webhook validator for Uri\WhatWg\Url, and in the same PR a different file used Uri\Rfc3986\Uri for our internal service identifiers. He read that as PHP 8.5 failing to make up its mind. I read it as the most honest thing the URL story in PHP has had since parse_url() arrived back in PHP 4. Having two classes is the fix. It isn't a leftover.
Quick recap in case you haven't touched 8.5 yet. The new URI extension is bundled, so there's nothing to pull in through Composer. Uri\Rfc3986\Uri is built on the uriparser library and follows RFC 3986, the general grammar for identifiers, which covers URNs like urn:uuid:..., custom schemes and relative references such as /path/foo. Uri\WhatWg\Url runs on Lexbor, the engine PHP 8.4 already brought in for the DOM API, and it reads URLs the same way a browser does. It won't take a relative reference unless you give it a base, and it knows that münchen.example and its Punycode form are one host, which is why you get getUnicodeHost(), getAsciiHost() and toAsciiString(). Both classes are immutable, both give you with*() methods in the PSR-7 style, and both have an equals() where you can pass UriComparisonMode::IncludeFragment if the fragment should count.
My reason for caring about the split is a particular kind of bug that has cost me a few nights. Your validator reads a URL and decides the host is allowed. Then your HTTP client, or curl underneath it, or a proxy further down, reads the same bytes and connects somewhere else. No single component is wrong. They just disagree about the string, and an SSRF lives in that disagreement. The PHP documentation says as much: running validation and resource fetching through different parsers can open security holes. parse_url() never tried to follow any standard, so it could disagree with pretty much everything, and for years we pretended that array of keys that may or may not be there counted as a parse.
With a single "Uri" class, the internals team would have had to pick a winner, and whichever standard lost would have been bent into the winner's shape. Remember the old joke that PHP has three ways to do everything? This is the other case: two ways because the world really does have two definitions, and the language decided to stop covering that up. If the string goes to an HTTP client, or came from a browser, or will be sent back as a redirect, use WHATWG, because the rest of the chain already reads it that way. If it's an identifier your own systems invented, such as a URN, a queue address or a my-app:// deep link, use RFC 3986, since a browser model has no business judging those.
The best argument against me deserves a proper hearing. Most of us never meant to become standards lawyers. A junior who needs the host out of a link now has to learn why two specs exist before writing one line, and the page explaining it is longer than the old parse_url() docs ever were. PSR-7 implementations have handled immutable URIs for years, and libraries like league/uri have covered the edge cases for a decade, so you could fairly ask what core adds besides a second vocabulary. On top of that, parse_url() is still there, still fast, and still fine for pulling the path out of a URL you built yourself thirty lines up.
I accept nearly all of that and still come down the same way, for two reasons. The first is that a userland library can't make your whole stack agree. A core class that framework HTTP clients, validators and PSR-7 adapters can all build on can, over time, and a shared parser is exactly what closes the mismatch gap. The second is the learning curve. Figuring out whether your string is a web link or an identifier is the real question, and parse_url() let you avoid it right up until a bug report made you answer it. Url::parse() returning null and filling an $errors array with UrlValidationError objects is also plainly nicer for form input than a try/catch around every field.
So here's my rule for our codebase, and feel free to steal it. Wherever a URL crosses a trust boundary (user input, webhooks, redirects, anything that ends up in an outbound request), parse_url() is banned, and the parser that validates has to be the same class the fetching code consumes. Everywhere else, move over when you're in that file anyway. Rewriting three hundred harmless call sites in one sprint just turns a sensible migration into a blame-log incident.
What I don't have figured out yet is the seam between the two. Your PSR-7 request object holds a URI, your outbound client probably wants WHATWG semantics, and your router probably thinks in RFC 3986 terms. Where in your application do you convert, and which class goes into your domain objects: one type everywhere, or both, with the boundary made explicit? I'd really like to see how your teams are drawing that line.




Comments
No comments yet — be the first.
Open the discussion
No account or password needed — just enter your e-mail and we’ll send you a one-time sign-in link. First time here? You’re set up automatically.
Your rating will be applied automatically after you sign in.
Check your inbox
We’ve sent a sign-in link to …. Open it on this device — this tab will sign you in automatically.
Nothing arrived? Check your spam folder — and mark the mail as "Not spam" so it lands in your inbox next time.