Uncategorized

How to Extract and Read HAR Files (Network Logs) Safely

How to Extract and Read HAR Files (Network Logs) Safely

To extract and read a HAR file safely, open the network tab in your browser’s developer tools (F12), replicate the issue to record network requests, right-click the log, and select “Save all as HAR with content”. You can then read the HAR file using a local, in-browser HAR file extractor to inspect headers, responses, and payload data without uploading sensitive session cookies to a remote server.

Key Definitions: HTTP Archive Format, JSON Structure, and Network Tokens

Understanding the precise terminology surrounding the HTTP Archive Format is critical for diagnosing web application performance anomalies and debugging complex API interactions safely and efficiently.

  • HTTP Archive Format (HAR): A JSON-formatted archival standard specifically utilized for logging a web browser’s interaction with a site. It contains highly detailed performance data, HTTP request and response headers, and the precise body content transferred during a browsing session, serving as a forensic snapshot of the network state.
  • W3C Web Performance Working Group: The international consortium that initially developed and maintained the HAR 1.2 specification to standardize how modern browsers export network telemetry, HTTP transactions, and performance metrics across different user agents.
  • JSON Structure: The hierarchical object notation used natively within a HAR file, typically starting with a root log object that encapsulates the creator, browser, pages, and highly detailed entries arrays. The deterministic nature of JSON allows for programmatic parsing and analysis.
  • Network Payloads: The actual data transmitted in the body of an HTTP request or response. This includes HTML documents, structured JSON API responses, binary image assets (frequently base64-encoded within the HAR), and multipart form submissions.
  • Session Tokens: Cryptographic identifiers (such as JSON Web Tokens (JWT), session cookies, and OAuth Bearer tokens) included in HTTP headers that authenticate the user. Because HAR files capture the exact headers, these tokens must be meticulously scrubbed before sharing a HAR file to prevent account compromise.

Why Developers and Support Teams Request HAR Files

When a modern single-page application (SPA) exhibits erratic behaviour, slow page load times, or silent API failures that cannot be reliably reproduced in an isolated testing environment, developers and customer support engineers will request HAR files. These network logs provide a granular, deterministic, and objective record of the exact HTTP transactions that occurred on the client’s machine at the exact time of failure.

By conducting a rigorous analysis of a HAR file, infrastructure engineers can identify stalled DNS lookups, blocked Cross-Origin Resource Sharing (CORS) preflight OPTIONS requests, misconfigured HTTP headers, and missing CDN assets resulting in 404 HTTP status codes. Furthermore, HAR files contain the exact API responses from backend microservices, allowing developers to see if the server returned malformed JSON, unexpected error payloads, or internal server errors (500s). This circumvents the ambiguity of vague user reports, transforming subjective complaints into actionable, empirical engineering data. It serves as an immutable snapshot of the client-server dialogue, making it the ultimate tool for resolving elusive edge cases and intermittent race conditions.

Step-by-Step: How to Generate a HAR File in Google Chrome (DevTools)

Google Chrome remains the industry-standard environment for web debugging, utilizing the robust Chrome DevTools suite to capture and export network traffic with high fidelity.

  1. Navigate to the specific web application or page where the error, latency, or issue occurs.
  2. Open the Chrome Developer Tools by pressing F12 on Windows/Linux or Cmd + Option + I on macOS. Alternatively, you can right-click anywhere on the rendered page and select Inspect from the context menu.
  3. Click on the Network tab located at the top of the DevTools panel to reveal the network logging interface.
  4. Ensure the circular recording button in the upper left corner of the Network tab is red. If it is grey, click it to begin recording network traffic immediately.
  5. Check the Preserve log checkbox. This is a critical step; it ensures that network requests are not unilaterally cleared when the page navigates, redirects, or reloads, capturing the full traversal path.
  6. Optional but recommended: Check the Disable cache box (while DevTools is open) to force the browser to request all assets fresh from the server, preventing 304 Not Modified responses from obscuring the true payload.
  7. Refresh the page and perform the exact sequence of actions required to reproduce the issue.
  8. Once the issue has been definitively captured, right-click anywhere within the grid of recorded network requests and select Save all as HAR with content.
  9. Save the generated .har file to a secure directory on your local workstation.

Step-by-Step: How to Export a HAR File in Mozilla Firefox

Mozilla Firefox provides equivalent, highly capable network inspection tools within its Developer Tools suite, adhering strictly to the standardized HAR specification for broad compatibility.

  1. Open the Mozilla Firefox browser and navigate to the web page exhibiting the problematic behaviour.
  2. Access the Developer Tools by pressing F12 or Cmd + Option + E on macOS, or select Web Developer > Network from the primary application menu.
  3. Click on the Network Monitor tab to bring the traffic capture grid into focus.
  4. In the upper right corner of the Network Monitor, click the gear icon to access the advanced settings dropdown and select Persist Logs. This guarantees data retention across page unloads and HTTP 301/302 redirects.
  5. Similar to Chrome, it is advisable to check Disable Cache to guarantee a complete payload fetch from the origin server.
  6. Reload the page and execute the precise steps necessary to trigger the software bug or performance bottleneck.
  7. After the network activity has completely ceased, right-click any populated row in the network request list.
  8. Select Save All As HAR from the contextual menu.
  9. Designate a secure local destination folder and save the archive file for subsequent analysis.

The Anatomy of a HAR File: log, creator, browser, pages, and entries arrays

A HAR file is fundamentally a massive JSON document structured according to a strict, predictable schema. Understanding this precise schema architecture is vital for parsing the data manually, building automated sanitization scripts, or extracting telemetry programmatically.

The root object is always designated as log, which encapsulates the entire network archive. Within this log object, several standardized arrays and configuration objects exist to provide context:

  • version and creator: Specifies the HAR specification version (typically 1.2) and identifies the specific software engine (e.g., Chrome, Firefox, Charles Proxy, Fiddler) that generated the telemetry log.
  • browser: Contains metadata regarding the browser name, version, and rendering engine utilized during the capture session.
  • pages: An array defining the individual HTML page loads captured in the session. It includes the pageTimings object (detailing metrics such as onContentLoad and onLoad) and a unique id for tracking each distinct page view.
  • entries: The core, most critical component of the HAR file. This massive array contains a discrete object for every single HTTP request initiated by the browser. Each entry includes:
    • startedDateTime: The exact ISO 8601 timestamp of the request initiation.
    • time: The total duration of the request lifecycle in milliseconds.
    • request: An object detailing the URL, HTTP method (GET, POST, PUT), headers, cookies, query string parameters, and postData.
    • response: An object detailing the HTTP status code, response headers, cookies, content payload (often base64 encoded for binary assets), and the accurately resolved mimeType.
    • timings: A highly granular breakdown of the request lifecycle, measuring DNS resolution time, SSL handshake latency, TCP connection time, Time to First Byte (TTFB), and content download time.

How to Extract Data from a HAR File in Your Browser (Viewing Requests & Responses)

Reading a raw, multi-megabyte HAR file in a standard text editor is highly inefficient and practically impossible due to its extensive, deeply nested JSON structure and massive base64-encoded strings. To analyse the data effectively, you must parse the file to visualize the specific HTTP transactions.

You can read and filter the file safely using a local, client-side HAR file extractor. Because these modern tools operate entirely within the DOM using secure JavaScript File APIs (specifically the FileReader API), your network log is never transmitted to an external server. Once the JSON is parsed into memory, you can aggressively filter the entries array by MIME type (e.g., isolating only application/json or application/graphql), search for specific HTTP status codes indicating failure (e.g., 400 Bad Request, 401 Unauthorized, 502 Bad Gateway), and inspect the raw request and response payloads. If your diagnostic task requires parsing specific nested values from complex JSON payloads found within the HAR responses, you can easily pipe the extracted body content into a dedicated utility to extract specific JSON fields.

Critical Security Warning: Why You Must Scrub Session Cookies and Authorization Headers

HAR files are designed to capture the exact, unadulterated state of your browser’s network requests, including all HTTP headers transmitted during the session. Consequently, they inherently contain the highly sensitive cryptographic credentials used to authenticate your user session.

If you record a HAR file while logged into an application (such as a banking portal, SaaS dashboard, or webmail client), the file will definitively contain your active session cookies (e.g., PHPSESSID, JSESSIONID, _session_id) and Authorization headers (e.g., Bearer eyJhbGci...). If a malicious actor intercepts or gains unauthorized access to this HAR file, they can extract these tokens and effortlessly perform a session hijacking attack. This grants them full, authenticated access to your account without requiring your password or bypassing your two-factor authentication (2FA) mechanisms. Before transmitting a HAR file to a third-party support team or attaching it to a public bug tracker, you must open the JSON file and meticulously scrub (redact, obfuscate, or delete) the values associated with the Cookie, Set-Cookie, and Authorization headers located within the request and response objects of the entries array.

HAR vs PCAP vs Log Files: What’s the Difference?

It is structurally essential to distinguish between different types of diagnostic telemetry files used in network engineering and software debugging.

  • HAR (HTTP Archive): Operates strictly at Layer 7 (Application Layer) of the OSI model. It is generated by the web browser and records decrypted HTTP, HTTPS, HTTP/2, and HTTP/3 requests and responses. The browser abstracts away the binary framing of HTTP/2 and presents a unified JSON structure. It is the ideal format for debugging web application logic, API contracts, and frontend asset loading.
  • PCAP (Packet Capture): Operates primarily at Layer 3 and Layer 4 (Network and Transport Layers). Generated by sophisticated network sniffers like Wireshark or tcpdump, a PCAP file captures raw TCP, UDP, and IP packets. If the traffic is encrypted via TLS (which governs modern HTTPS), a PCAP file cannot expose the HTTP payloads without the corresponding, highly restricted SSL/TLS pre-master secret session keys. PCAP files are strictly used for diagnosing low-level network routing anomalies, TCP retransmissions, packet loss, and complex TCP handshake issues.
  • Server Log Files: Generated by backend web servers (e.g., Nginx, Apache HTTP Server) or backend application frameworks. These logs record incoming requests exclusively from the server’s perspective. While they reliably log request URIs, IP addresses, user agents, and HTTP status codes, they rarely log the full request and response body payloads in order to conserve disk I/O and storage space. This limitation makes client-side HAR files vastly superior for front-end debugging.

Privacy & Security: Why You Should Never Upload HAR Files to Third-Party Online Viewers

The internet currently hosts numerous “free HAR viewer” websites that ostensibly offer convenience by requiring you to upload your `.har` file to their remote servers for parsing, indexing, and visualization.

Uploading a HAR file to a remote, unverified server poses a catastrophic security and privacy risk. You are essentially transmitting your active session tokens, proprietary API responses, Personally Identifiable Information (PII) protected under GDPR and CCPA, internal network architecture details, and potentially intellectual property directly to an untrusted third party. Even if the service explicitly claims to delete the file immediately after parsing, the data may be inadvertently logged in their server access logs, cached in load balancers, or intercepted during transit if TLS is improperly configured. Always utilize local, client-side extraction utilities that utilize Web Workers and DOM parsing to keep the JSON file strictly within your local browser memory, guaranteeing absolute zero data exfiltration.

Frequently Asked Questions

What program opens a HAR file?

A HAR file is, at its core, a standardized JSON text file that can be natively opened with any advanced text editor (such as Visual Studio Code, Sublime Text, or Notepad++). However, for practical, human-readable analysis, it must be parsed using a dedicated HAR viewer application or a local browser-based extractor. These tools ingest the JSON arrays and render the requests and responses in a structured, sortable data grid, mimicking the browser’s own Network tab interface.

Do HAR files contain passwords?

Yes, absolutely. If you submit a login form or authentication request while actively recording a network log, the HAR file will intercept and record the raw HTTP POST request. This record definitively includes the plaintext password or securely hashed credentials directly within the postData or request body payload. You must identify and redact this specific payload object before sharing the file with anyone.

How do I open a HAR file in Chrome?

You can seamlessly import an existing HAR file back into Google Chrome for analysis. Open the Chrome DevTools interface (F12), navigate to the Network tab, and simply drag and drop the `.har` file from your desktop directly into the network request grid. Chrome’s DevTools engine will parse the JSON and display the historical network data exactly as if it were recorded live in that session.

Can HAR files be forged or edited?

Yes, they are highly susceptible to modification. Because a HAR file is merely a standard, unencrypted JSON text document without cryptographic signatures or checksums, any user can open it in a text editor and manually modify timestamps, alter HTTP status codes, inject fake header values, or rewrite payload contents. Therefore, while HAR files are invaluable for cooperative software debugging, they are fundamentally unreliable as definitive legal evidence or irrefutable forensic logs.

Why is my HAR file so large?

HAR files routinely grow to exceptional sizes (often exceeding 50 to 100 megabytes) because they are designed to encode the complete, uncompressed response bodies of all loaded assets. This includes bulky base64-encoded representations of high-resolution images, massive minified JavaScript application bundles, comprehensive CSS stylesheets, and extensive JSON API responses. Every single byte transferred over the wire is logged into the JSON structure.

How do I reduce the size of a HAR file before sending?

To aggressively reduce the file size for easier transmission via email or ticketing systems, you can manually open the JSON and delete entries objects related to heavy, static assets (like `.png`, `.jpg`, `.css`, and `.woff2` font files), keeping only the relevant application/json XHR/Fetch API requests. Furthermore, because HAR files are purely text-based JSON, they compress exceptionally efficiently. Compressing the `.har` file into a `.zip` or `.gz` archive can often reduce its footprint by up to 80%.

What does HTTP Archive Format mean?

The HTTP Archive Format (HAR) is a highly standardized, JSON-based file format originally drafted and designed by the W3C Web Performance Working Group. Its explicit purpose is to provide a universal mechanism to export, share, and archive the detailed records of a web browser’s HTTP transactions with web servers, enabling cross-platform performance analysis and deterministic debugging.

Sources & Standards

This technical guide strictly adheres to the definitions, architectural constraints, and structural JSON schemas defined by the W3C Web Performance Working Group. For comprehensive programmatic implementation details, strict schema validation, and historical context, consult the official W3C HAR 1.2 Specification.

Keep reading