← All guides

Developers · 7 min read · Updated October 5, 2026

URL Encoding Explained: When to Percent-Encode a Link

A short link is a promise that the long URL on the other end still works. That promise breaks more often than people expect, and one of the quietest ways it breaks is encoding: the destination you pasted contained a character the URL format cannot carry raw, and either it never got encoded at all, or it got encoded twice by two tools that each assumed the other had not touched it yet.

This guide is the explanation most pages skip over: what percent-encoding actually does, the real difference between encoding a whole URL and encoding one value inside it, the two bugs that account for nearly every encoding support ticket, and where this shows up when you are building destination URLs, tagging campaigns, or calling the API directly.

What percent-encoding actually does

URLs are restricted by spec to a small set of ASCII characters: letters, digits, a few punctuation marks (-, ., _, ~), and the structural characters that give a URL its shape (: / ? # [ ] @ ! $ & ' ( ) * + , ; =). Everything else, spaces, accented letters, emoji, and reserved characters used outside their structural role, has to be represented instead of written literally. Percent-encoding does that: it takes a character, converts it to bytes using UTF-8, and writes each byte as a percent sign followed by two hexadecimal digits.

Common characters and the percent-encoded form a URL carries instead.
CharacterPercent-encodedWhy it needs encoding
space%20Not a legal URL character at all
&%26Reserved: separates query parameters
=%3DReserved: separates a key from its value
+%2BWould otherwise be read as an encoded space in form data
/%2FReserved: separates path segments
é%C3%A9Non-ASCII characters encode as multiple UTF-8 bytes

That last row is the one people miss: encoding is not one swap per character, it is one swap per byte of the UTF-8 representation, which is why a single accented letter or emoji can expand into two or three %XX triplets. Reversing the process, turning those triplets back into readable text, is URL decoding.

Component encoding vs full-URL encoding

There are two encoding modes, and picking the wrong one is the first way this goes wrong. Component encoding treats its input as one opaque value and escapes everything that is structurally significant in a URL, including ? & = /. Full-URL encoding assumes its input is already a complete, valid URL and leaves those structural characters alone, only escaping what would be illegal anywhere in a URL.

encodeURIComponent('a&b=c?')
// → 'a%26b%3Dc%3F'      (everything structural gets escaped)

encodeURI('https://example.com/search?q=a&b=c')
// → 'https://example.com/search?q=a&b=c'   (unchanged: & and = are
//    doing their normal job here, not sitting inside a value)

The rule that keeps these straight: encode with the component function when you have one value that is going to be inserted into a URL (a query parameter, a path segment), and reach for full-URL encoding, or no encoding at all, only when you already hold a complete, structurally correct URL. Mixing them up either leaves a dangerous value un-escaped or mangles a perfectly good URL's own structure.

The classic bug: a URL inside a URL

This comes up constantly with short links: a redirect that needs to carry a return path, an SSO callback, or a share link that embeds a deep link as one of its own parameters. The destination becomes a value inside another URL, and if it is not component-encoded first, its own ? and & get parsed as if they belonged to the outer URL.

# Broken: the inner URL's own query string leaks into the outer one
https://go.acme.com/r?next=https://example.com/path?ref=newsletter&utm_source=x
# A parser sees two params here: next=https://example.com/path?ref=newsletter
# and utm_source=x — not one value carrying the full inner URL.

# Fixed: component-encode the inner URL before it becomes a value
https://go.acme.com/r?next=https%3A%2F%2Fexample.com%2Fpath%3Fref%3Dnewsletter%26utm_source%3Dx

The fix is always the same: run the inner URL through component encoding before concatenating it into the outer one. Whatever reads the next parameter later decodes it once and gets the original URL back intact, utm_source and all.

Double-encoding: the other classic bug

The opposite mistake is encoding something that is already encoded. A literal % is itself not a legal raw character, so encoding one turns it into %25. Run that through a second time and %20 becomes %2520, a space becomes a percent sign followed by digits that spell out, once decoded, a different percent sign followed by more digits. The symptom is almost always one of two things: a decoded value that still shows %20 or %3D instead of the character it should be, or a destination that resolves to a 404 because the server received the literal string %2520 instead of a space.

Double-encoding usually happens in a pipeline, not in one place: a browser form already encodes what you typed, and code further downstream encodes it again before building a redirect or an API call, each step unaware the other already did the job.

To catch it, paste the suspect value into a URL parser and look for %25 followed by what is clearly another escape sequence, that is the tell. The fix is to decode until the value stops changing, confirm you are looking at the readable original, and then encode exactly once on the way back out.

Plus signs, spaces, and the form-encoding exception

One context breaks the %20-for-space rule on purpose. The application/x-www-form-urlencoded format, used by HTML form submissions and by convention in many search query strings, encodes a space as + instead of %20. Everywhere else, in a URL path, in standard percent-encoding, a space is %20 and a literal + is just a plus sign.

This is a frequent, silent source of breakage because email addresses and UTM values commonly contain a literal +, as in [email protected]. Decode that string with a form decoder, the kind many server frameworks default to for query strings, and the + turns into a space, silently corrupting the address. Decode it with a standard URL decoder and the + is left alone, exactly as typed. Knowing which decoder a given piece of code is using is the only way to predict which behavior you get.

  • Redirects that carry a return path. An auth callback or a campaign landing page that forwards a deep link needs its inner URL component-encoded, exactly as above, or the forward silently truncates at the first &.
  • UTM values with spaces or punctuation. A campaign name like "Black Friday" or "Q4 push" needs its space encoded to survive in a query string; see UTM naming conventions and what UTM parameters actually are for the naming side of this.
  • Links built programmatically against the API. Constructing a destination URL by string concatenation instead of a query-string library is exactly how double-encoding and unescaped & creep into API-created links.
  • Bulk-imported destination URLs. A spreadsheet of links from several sources often mixes already-encoded and raw values. Importing it as-is bakes in that inconsistency, and re-encoding everything blindly double-encodes the rows that were already correct.

Debugging an encoding problem in under a minute

  1. Paste the suspect URL into a URL parser and read the query string field by field. A value containing a stray %, &, or = in the wrong place is the tell that something inside it was never encoded, or was encoded wrong.
  2. Take the one value that looks off and paste it alone into a URL encoder/decoder, then decode it. If the result still contains %XX sequences, it was double-encoded, decode again until it stops changing.
  3. Starting from that clean, readable value, re-encode it exactly once with the correct mode, component for a single value, full-URL only if what you have is already a complete URL, then reassemble.
  4. If the problem is the query string as a whole rather than one value, build it with a query string builder instead of hand concatenation. It encodes each value in the right mode automatically, which removes the chance of mixing modes by hand.

What a well-formed destination URL buys you

A link form that validates with the browser's own URL parser, which is what ReSlug's destination field does, catches a malformed URL immediately: unbalanced brackets, a missing scheme, a stray raw space in the host. What it cannot catch is a meaning-level mistake like a value that is double-encoded but still happens to parse as a technically valid URL, since nothing about that string is actually broken, it is just pointing somewhere other than where you meant. That class of bug only shows up once a click lands on the wrong page or a 404, which is why it is worth checking the destination with a parser before the link goes out, not after someone reports it.

Frequently asked questions

What is URL encoding (percent-encoding)?

It is the way a URL represents a character it cannot carry literally, a space, an accented letter, a reserved character like & used outside its structural role. The character is converted to UTF-8 bytes, and each byte is written as a percent sign followed by two hexadecimal digits, for example %20 for a space.

When do I actually need to encode a URL myself?

Only when you are assembling a URL from parts: inserting a value that might contain reserved characters into a query string or path, or embedding one complete URL as a parameter of another. Clicking an existing link, or pasting a normal web address into a browser, never requires manual encoding.

What is the real difference between encodeURIComponent and encodeURI?

encodeURIComponent escapes everything structurally significant, including ? & = and /, so it is correct for a single value headed into a URL. encodeURI assumes the input is already a complete, valid URL and leaves those structural characters alone, so using it on a single value, instead of a full URL, leaves reserved characters dangerously un-escaped.

Why does decoding my URL still show %20 or %3D instead of the actual character?

The value was encoded more than once. The first encoding pass is correct, but a second pass turns the percent sign from that first pass into %25, so %20 becomes %2520. Decoding it once only partly undoes the damage; decode repeatedly until the string stops changing to reach the original.

Is a plus sign the same thing as %20?

Only inside application/x-www-form-urlencoded data, the format HTML forms submit with, where + is the specific encoding for a space. In a URL path, and in standard percent-encoding generally, a space is %20 and a literal plus sign stays a plus sign, which matters a great deal for values like [email protected].

Keep reading