Base64 size overhead:
why encoded data is about a third bigger
Predict encoded size before you embed an image, send an attachment or hit a payload limit, and know what the = signs mean.
Calcylator Editorial Team
Updated · 5 min read
Three bytes in, four characters out
Base64 exists to carry arbitrary binary data through channels designed for text, such as JSON fields, e-mail bodies and URLs. It does this by taking the input three bytes (24 bits) at a time and splitting those bits into four groups of six. Each six-bit group maps to one of 64 safe characters.
So three bytes always become four characters. For example the ASCII text Man, three bytes, encodes to TWFu. The expansion is 4 ÷ 3, which is a 33.3 percent increase, and it is a property of the method, not of the data.
- n:
- input size in bytes
- ⌈ ⌉:
- round up to the next whole number
Input
1,000,000 bytes
Encoded size
1,333,336 characters (+33.3%)
1,000,000 ÷ 3 = 333,333.33, rounded up to 333,334; × 4 = 1,333,336.
What the = padding does
When the input length is not a multiple of three, the last group is short. Base64 fills the output to a whole four characters using = signs so decoders can tell how many real bytes the final group held.
| Input bytes | Output characters | Padding |
|---|---|---|
| 1 | 4 | == |
| 2 | 4 | = |
| 3 | 4 | none |
| 4 | 8 | == |
| 5 | 8 | = |
| 6 | 8 | none |
Padding is why tiny inputs look disproportionately large: a single byte costs four characters. Over big payloads it is at most two characters and vanishes in the overall percentage.
Encoded sizes at a glance
A quick mental rule is to add one third and round up a little. The exact values from 4 × ceil(n ÷ 3) are below.
| Original size | Base64 length | Growth |
|---|---|---|
| 10,000 bytes | 13,336 | +33.4% |
| 100,000 bytes | 133,336 | +33.3% |
| 1,000,000 bytes | 1,333,336 | +33.3% |
| 5,000,000 bytes | 6,666,668 | +33.3% |
Notice that the percentage settles at one third almost immediately. Only for very short inputs, a few bytes to a few dozen, does padding visibly change the ratio.
If the encoded text is then placed in a URL query, standard Base64 characters such as + and / and the = padding are percent-encoded as three characters each, which can inflate the size again. That is the reason the URL-safe alphabet, using − and _, exists.
Line breaks, data URIs and other extras
Standards for e-mail (MIME) wrap Base64 at 76 characters per line, adding a carriage return and line feed to each. That inflates the size a bit more. The 1,333,336 characters above would split into 17,544 lines, adding about 35,000 bytes for a total of 1,368,424, which is 36.8 percent over the original.
- Data URIs add a prefix such as data:image/png;base64, of a couple of dozen characters.
- JSON strings may escape some characters, though Base64 characters need no escaping.
- If the text is then stored as UTF-16, each character takes two bytes, so the stored size doubles again.
What the overhead means for real payloads
A 300 KB image becomes 400 KB as Base64. Embedding it in a page via a data URI saves a request but makes the HTML bigger, blocks parsing, and prevents separate caching, so it only pays off for small icons.
API payload limits are another place it bites. A service that accepts 1 MB request bodies fits at most about 750 KB of raw file once Base64 is used, since 1,000,000 ÷ 4 × 3 = 750,000. Mail gateways with a 25 MB limit similarly allow attachments of roughly 18 to 19 MB.
Gzip or Brotli on the transport recovers a lot of this for compressible content, but compressed formats such as JPEG will not shrink much, so the 33 percent largely remains.
Alternatives and when Base64 is still right
Base64 is not encryption or compression. It is a representation. Anyone can decode it, and it always makes the data larger.
When a binary-safe channel is available, such as multipart form uploads, a binary protocol or object storage with a signed URL, send the bytes directly and skip the overhead. Base64 remains the sensible choice for small secrets, signatures, cryptographic keys, inline icons and any place where only text can travel. Before choosing, run the numbers with a Base64 size calculator and see whether the extra third is acceptable.
The decoding side
A decoder reverses the mapping, and the length tells it a lot. A valid padded string has a length that is a multiple of four. If the last group ends in one = sign, the final group held two bytes, and with == it held one.
- Original size ≈ length × 3 ÷ 4, minus the number of = signs.
- A length that leaves a remainder of 1 after dividing by 4 is never valid Base64.
- Line breaks and whitespace are ignored by most decoders, but should be stripped before counting.
Decoding in memory needs space for both the text and the result, which means roughly 1.75 times the original size at once. For large files, decode in a stream rather than loading a huge string, or avoid Base64 for the transfer entirely.
Variants and how they change length
Several flavours share the same 3-to-4 mapping but differ in details that affect size and compatibility.
| Variant | Alphabet and rules | Length effect |
|---|---|---|
| Standard (RFC 4648) | A–Z, a–z, 0–9, + and /, padded with = | 4 × ceil(n ÷ 3) |
| URL-safe | Uses − and _ in place of + and /; padding is often dropped | Up to 2 characters shorter |
| MIME | Standard alphabet, lines of 76 characters | About 2.6% longer than standard |
Pick the one the receiving system expects, and be consistent. A token encoded with the URL-safe alphabet will fail a strict standard decoder, and the reverse, so the size is the lesser concern next to compatibility.
Lastly, do not forget the cost of the format at the other end of the pipe. Encoding and decoding take CPU time and temporary memory, and logs that capture request bodies will store the inflated text. For one-off small values this is irrelevant, but for a service that handles thousands of uploads per minute, shifting to a binary upload path can pay back the engineering effort quickly.
Common questions
How much bigger is Base64 than the original?
About 33 percent bigger. Every three input bytes become four characters, so the output length is 4 × ceil(n ÷ 3). A 300 KB file becomes about 400 KB, plus a few padding characters at the end.
What is the formula for Base64 encoded length?
Encoded length equals 4 times the ceiling of n divided by 3, where n is the input size in bytes. One byte gives 4 characters, 3 bytes give 4, and 4 bytes give 8, with = padding filling the last group.
Why does Base64 end with = signs?
The = signs pad the final group to four characters when the input is not a multiple of three bytes. One missing byte gives one =, two missing give ==, and an exact multiple gives none.
How do I get the original size from a Base64 length?
Multiply the character count by 3 ÷ 4, then subtract one byte for each = sign at the end. For 1,333,336 characters ending in ==, that is 1,000,002 − 2 = 1,000,000 bytes.
Does Base64 encrypt or compress data?
Neither. It only converts bytes to printable text so they travel safely, and anyone can decode it. It always increases size by about a third, so do not rely on it for secrecy or efficiency.
Was this guide helpful?
Continue reading
View all blogsPagination Math: How Many Pages and Offsets?
Pages = ceil(total ÷ page size). 1,037 records at 25 per page is 42 pages, and the last holds just 12. Offsets and edge cases explained.
5 min read
Binary to Decimal Conversion, Step by Step
Convert binary to decimal by adding the place values of each 1: 101101 is 32 + 8 + 4 + 1 = 45. Includes the doubling shortcut and an IP address example.
5 min read
CIDR Ranges Explained: Mask, Hosts and Block Size
A /26 holds 64 addresses and 62 usable hosts. See how 2^(32−n) works, how to find the network and broadcast address, and how to size a subnet.
5 min read




