When writing backend APIs or debugging frontend analytics, you often need to programmatically extract variables from a URL. Trying to parse these strings manually with split("&") or regular expressions is a recipe for disaster when you encounter URL-encoded characters or duplicate keys.

This guide breaks down the correct, native APIs to read URL query strings programmatically across JavaScript, Python, and Bash, explores the specific edge cases of tracking UTM parameters, and details the real-world security risks of query string leaks.

The RFC 3986 Query Component

According to RFC 3986, a URI query component is indicated by the first question mark (?) character and terminated by a number sign (#) or by the end of the URI.

It typically consists of key-value pairs separated by ampersands (&): https://example.com/login?utm_source=newsletter&user_id=12345

For a visual breakdown of a full URL into its respective hostname, path, and search parameters, you can use a URL Parser.

Extracting UTM Parameters for Analytics

Urchin Tracking Module (UTM) parameters are a specific convention used by marketers to track campaign performance. There are five classic UTM keys (though GA4 also supports modern additions like utm_id, utm_source_platform, and utm_creative_format):

  1. utm_source (e.g., google, newsletter)
  2. utm_medium (e.g., cpc, email)
  3. utm_campaign (e.g., spring_sale)
  4. utm_term (Paid keywords)
  5. utm_content (A/B testing ad variants)

Key Gotchas when parsing UTMs:

  • Case-Sensitivity: utm_source=Google and utm_source=google will appear as two entirely different traffic sources in your analytics platform unless you explicitly .toLowerCase() them during ingestion.
  • Leaking via Canonical/Shared URLs: If a user copies a link from your site with their UTM parameters still attached, they will artificially inflate that campaign if they share it on Reddit. You should dynamically strip UTM parameters using the History API (history.replaceState()) after your analytics tag has recorded the page view, otherwise GA4 will lose the campaign data entirely. Additionally, ensure your <link rel="canonical"> tag is set to the clean URL without UTMs.
  • GA4 Logging: By default, Google Analytics 4 records the entire URL into the page_location dimension, meaning any sensitive query parameters living alongside the UTMs will be stored on Google’s servers.

How to Read Query Strings in JavaScript

Never use split('&') or regex to parse URLs in JavaScript. The modern, robust approach is the native URLSearchParams interface (MDN Reference), which automatically handles URL decoding.

// Assume current URL is: https://site.com/?product=shoes&color=blue%20metallic

const queryString = window.location.search;
const urlParams = new URLSearchParams(queryString);

// Extract specific values
const product = urlParams.get('product'); // "shoes"
const color = urlParams.get('color'); // "blue metallic" (automatically decoded)

// Handling duplicate keys / arrays (e.g., ?arr[]=1&arr[]=2)
const arrays = urlParams.getAll('arr[]'); // ["1", "2"]

How to Read Query Strings in Python

In Python, the standard library provides everything you need in the urllib.parse module (Python Docs).

from urllib.parse import urlparse, parse_qs

url = "https://dashboard.app.com/reports?user=admin&token=abc123xyz"

parsed_url = urlparse(url)

# Convert the query string into a dictionary
# keep_blank_values=True ensures empty keys like ?a=&b=2 don't disappear
params = parse_qs(parsed_url.query, keep_blank_values=True)

# Note: parse_qs returns a list for each value to handle duplicate keys
print(params) 
# Output: {'user': ['admin'], 'token': ['abc123xyz']}

token = params.get('token', [None])[0]

How to Read Query Strings in the CLI / Bash

When parsing logs or curling endpoints, you might need to extract parameters directly from the terminal. Because regex frequently drops parameters when URLs contain multiple ? characters or encoded strings, the most robust one-liner leverages Python’s parser inline:

#!/bin/bash
URL="https://example.com/api?next=/a?b=1&c=2&q=blue%20metallic"

python3 -c 'import sys,urllib.parse as u;[print(f"{k}={v}") for k,v in u.parse_qsl(u.urlsplit(sys.argv[1]).query, keep_blank_values=True)]' "$URL"

Output:

next=/a?b=1
c=2
q=blue metallic

If you absolutely must use native sed without Python installed, this snippet will cut the string at the first ? (but leaves the values strictly encoded):

echo "$URL" | sed -n 's/^[^?]*?\([^#]*\).*/\1/p' | tr '&' '\n'

The 4 Query String Parsing Gotchas

When writing your own logic instead of relying on a dedicated tool like our Query String Parser, you will inevitably encounter these four edge cases:

Edge CaseThe ProblemThe Solution
Duplicate KeysCalling .get('id') on ?id=1&id=2 only returns "1".Use .getAll('id') in JS, or access the full array returned by Python’s parse_qs.
Plus Signs (+)?q=a+b decodes to "a b" (space). ?q=a%2Bb decodes to "a+b".Native parsers handle this automatically (unlike decodeURIComponent('a+b') which incorrectly returns "a+b"). Do not use generic string replacement.
Blank Values?user=&active=true will silently drop the user key entirely in Python.Always pass keep_blank_values=True to parse_qs().
Array Notation?arr[]=1&arr[]=2 is a PHP/Rails convention. Browsers do not automatically group these into a JS array.You must explicitly query the literal key including brackets: urlParams.getAll('arr[]').

The Security Risk: Query String Leaks

It is a common anti-pattern to place sensitive information—like password reset codes (?token=xyz), email addresses, or API keys—directly into URL query strings.

While HTTPS encrypts the URL during transit, query strings are notoriously leaky at rest.

How Query Strings Leak:

  1. HTTP Referer Headers: While modern browsers default to strict-origin-when-cross-origin (sending only the domain cross-site), the full URL still leaks same-origin, to third-party scripts on the page reading location.href, or if a site sets a looser policy like unsafe-url.
  2. Browser History: The full URL remains permanently logged in the user’s browser history, visible to anyone who has access to their machine.
  3. Analytics Tools: Tools like Google Analytics inherently track the full page_location. Unless explicitly redacted, you will upload sensitive PII directly into your analytics database. (Tip: Use GA4’s built-in Data redaction setting under Admin → Data streams → Redact data to automatically filter emails and chosen query parameters).
  4. Access Logs & CDNs: Load balancers, NGINX access logs, and CDNs (like Cloudflare or CloudFront) log full request URLs by default.

For API calls, always transmit sensitive tokens inside the HTTP POST Body or via Authorization headers. For magic links and password resets that must use URL tokens, ensure they are short-lived and single-use. Exchange them immediately, redirect the user to a clean URL, enforce Referrer-Policy: no-referrer on the landing page, and avoid loading third-party scripts on that specific route.