Documentation

This page is the map of how CompareStack behaves. Processing details that come from the code live on How CompareStack works. Worked inputs live under Examples. Longer workflow notes are the guides.

CompareStack is maintained by Tarak Moyyi. To report a wrong statement or a result that does not match this page, use Contact and include the page URL.

Getting started

What CompareStack is

Nine browser forms backed by server-side processing: four comparison tools, three formatters, and two encoders. There is no account and no saved history. The useful part of the site, besides the forms, is the written limit of each form.

Choosing a tool

File processing and privacy

Nothing runs until you submit the form. Paste fields accept up to 100,000 characters. Each upload is limited to 10 MB. File tools read PHP’s temporary upload and do not move it into application storage. Paste tools send the text in the POST body. Neither path is a private vault: server logs can still contain IP address, URL, and, on an extraction or Excel failure, the exception message. Read what happens when you upload a file and the Privacy Policy before you send regulated data.

Upload limitations

  • PDF must be PDF. Word must be DOC or DOCX. Spreadsheets must be XLS, XLSX, or CSV.
  • Extracted text over about 3 MB is rejected.
  • Excel stops each sheet after 1,200 data rows.
  • Paste endpoints are limited to 30 requests per minute per IP. Upload endpoints are limited to 10.

Document comparison

Text comparison

Lines are split on any newline. Matching lines stay unmarked. A removed line followed by an added line is a changed pair, with inline highlights split on whitespace. CRLF versus LF alone does not mark every line. Trailing spaces do. The tool does not understand that a number is a timeout or that a flag enables a cache.

PDF comparison

PDF Compare reads the text layer, rejoins wrapped lines, and splits sentences on . ! ? ;. It does not compare pixels and it does not run OCR. A scanned page with no selectable text comes back empty or failed. Details: how PDF text extraction works.

Word comparison

Word Compare collects paragraph and table text. If that pass is empty, DOCX is read from word/document.xml inside the package. Track Changes authors, comments, and formatting are not loaded. The same words in a different style are not a diff.

Excel comparison

Rows stay rows. Cells are joined with | . Empty cells become ∅. You see that a row’s values changed. You do not get a cell address, a formula audit, a chart diff, or conditional formatting.

Developer utilities

JSON formatting

The server decodes JSON and pretty-prints it when the parse succeeds. Trailing commas, comments, and single quotes are rejected. The tool does not apply JSON Schema and it does not repair the document. See the trailing-comma example.

SQL formatting

The SQL formatter inserts line breaks and uppercases a fixed keyword list. It is not a dialect parser and it does not change how a database executes the query. The compressed-query example shows the real output, including the split between INNER and JOIN.

Code formatting

Non-HTML input gets brace-and-bracket indentation. It is not Prettier, Black, or gofmt. The indentation example is the actual pass on a one-line function.

Encoding

URL encoding

Encode is rawurlencode (spaces become %20). Decode is rawurldecode. The whole paste is encoded, including :, /, ?, and & if you paste a full URL. Encode the value when you only need the query component changed. See the space and apostrophe example.

Base64 encoding

Standard Base64 uses +, /, and = padding. It is not encryption and it is not a password hash. URL-safe Base64 is not converted for you. See encoding is not encryption.

Troubleshooting

The PDF has no text

Open the file and try to select a sentence. If you cannot, the file is probably a scan or an image-only page. This site does not OCR it. Run OCR in a tool that does, then compare the resulting text, or compare a text-based export.

Why do all lines appear different?

On Text Compare, line endings alone are not the cause. Look for trailing spaces, a byte-order mark, non-breaking spaces, or a reflow that put different words on each line. On PDF Compare, a header or page number in the text layer can repeat as a change on every page, and sentence splitting can make a wrapped paragraph look like several edits. Compare a copied paragraph in Text Compare when you need to separate extraction noise from a real edit.

Invalid JSON

The error is the PHP parser message, prefixed with “Invalid JSON input:”. A trailing comma is “Syntax error”. Remove the comma, comments, and single quotes, then submit again. The formatter will not fix the payload for you.

Large file comparison

Over 10 MB, the upload is rejected. Over about 3 MB of extracted text, the comparison is rejected. Excel sheets stop at 1,200 data rows. Very large line diffs (more than 250,000 line-pairs in the alignment grid) use a simpler fallback instead of the full alignment. Split the input when you hit those limits.

Excel comparison limitations

Same displayed numbers can hide a formula that was replaced by a constant. Charts, colors, comments, and hidden sheets you did not expect to matter are outside the cell-value extract. Use Excel when the change is in the workbook structure rather than in the values.

Reporting a correction

If a page describes behavior the tool does not have, email support@comparestack.in from the contact page. Include the URL and what you observed. Documented corrections are listed on the changelog.