<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
    <title>Snacks - float</title>
    <subtitle>Snack-size learning for a fast-paced world.</subtitle>
    <link rel="self" type="application/atom+xml" href="https://snacks.devtestonly.uk/tags/float/atom.xml"/>
    <link rel="alternate" type="text/html" href="https://snacks.devtestonly.uk"/>
    <generator uri="https://www.getzola.org/">Zola</generator>
    <updated>2026-10-06T04:37:33+00:00</updated>
    <id>https://snacks.devtestonly.uk/tags/float/atom.xml</id>
    <entry xml:lang="en">
        <title>`float` and `double`</title>
        <published>2026-10-06T04:37:33+00:00</published>
        <updated>2026-10-06T04:37:33+00:00</updated>
        
        <author>
          <name>someone</name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://snacks.devtestonly.uk/c/float-and-double/"/>
        <id>https://snacks.devtestonly.uk/c/float-and-double/</id>
        
        <content type="html" xml:base="https://snacks.devtestonly.uk/c/float-and-double/">&lt;h2 id=&quot;float-and-double-comparison&quot;&gt;&lt;code&gt;float&lt;/code&gt; and &lt;code&gt;double&lt;/code&gt; comparison&lt;/h2&gt;
&lt;p&gt;Here is how a double data type directly compares to a float data type under the standard IEEE 754 specifications:&lt;/p&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Feature&lt;/th&gt;&lt;th&gt;float (Single Precision)&lt;/th&gt;&lt;th&gt;double (Double Precision)&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;Size in Memory&lt;/td&gt;&lt;td&gt;4 bytes (32 bits)&lt;/td&gt;&lt;td&gt;8 bytes (64 bits)&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;Sign Bit&lt;/td&gt;&lt;td&gt;1 bit&lt;/td&gt;&lt;td&gt;1 bit&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;Exponent Width&lt;/td&gt;&lt;td&gt;8 bits&lt;/td&gt;&lt;td&gt;11 bits&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;Mantissa Width&lt;/td&gt;&lt;td&gt;23 bits (24 effective)&lt;/td&gt;&lt;td&gt;52 bits (53 effective)&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;Decimal Precision&lt;/td&gt;&lt;td&gt;~7 digits of accuracy&lt;/td&gt;&lt;td&gt;~15 to 17 digits of accuracy&lt;/td&gt;&lt;/tr&gt;
&lt;tr&gt;&lt;td&gt;Approximate Range&lt;/td&gt;&lt;td&gt;±1.4 × 10⁻⁴⁵ to ±3.4 × 10³⁸&lt;/td&gt;&lt;td&gt;±5.0 × 10⁻³²⁴ to ±1.7 × 10³⁰⁸&lt;/td&gt;&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;h3 id=&quot;key-differences-explained&quot;&gt;Key Differences Explained&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Precision&lt;/strong&gt;: A double provides more than twice the decimal precision of a float. If you calculate &lt;code&gt;1.0 / 3.0&lt;/code&gt;, a float cuts off around &lt;code&gt;0.3333333&lt;/code&gt;, while a double carries out to roughly &lt;code&gt;0.3333333333333333&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Range&lt;/strong&gt;: A double can handle drastically larger (and smaller fractionally) numbers because its exponent component has 3 extra bits, allowing it to scale up to &lt;code&gt;10³⁰⁸&lt;/code&gt; compared to the float’s limit of &lt;code&gt;10³⁸&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Performance &amp;amp; Memory&lt;/strong&gt;: A float takes up half the memory space. In applications processing millions of numbers (like graphics programming or audio buffers), using float saves substantial memory and can run faster on hardware optimized for it.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr /&gt;
&lt;h2 id=&quot;how-are-the-mantissa-bits-used-in-float&quot;&gt;How are the mantissa bits used in float?&lt;/h2&gt;
&lt;p&gt;In a float (32-bit), the 23 mantissa bits (also called the significand or fraction) are used to store the actual precision digits of a number in binary scientific notation.
The IEEE 754 standard uses a clever trick called the implicit leading bit to get 24 bits of precision out of only 23 bits of physical storage. Here is exactly how it works:&lt;/p&gt;
&lt;h3 id=&quot;1-the-normal-form-the-implicit-1-trick&quot;&gt;1. The Normal Form: The Implicit “1.” Trick&lt;/h3&gt;
&lt;p&gt;In standard base-10 scientific notation, you always format numbers so there is exactly one non-zero digit before the decimal point (e.g., 4.51 × 10³).
Binary works the same way. In binary, the only non-zero digit is 1. Therefore, every normalized binary scientific number always looks like:
$$1.xxxxxxxxxxxxxxxxxxxxxxx_2 \times 2^{\text{exponent}}$$
Because the digit before the binary point is always 1, the hardware engineers realized they don’t actually need to waste memory storing it.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The 23 bits in memory only store the fractional part (the digits after the binary point).&lt;/li&gt;
&lt;li&gt;When the CPU performs calculations, it automatically re-attaches the 1. to the front.&lt;/li&gt;
&lt;li&gt;This gives you 24 bits of precision using only 23 bits of space.&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id=&quot;visual-example&quot;&gt;Visual Example:&lt;/h4&gt;
&lt;p&gt;If you want to store the number 9.0 in a float:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;9.0 in binary is 1001.0&lt;/li&gt;
&lt;li&gt;Move the binary point to normalize it: 1.001000… × 2³&lt;/li&gt;
&lt;li&gt;Drop the leading 1..&lt;/li&gt;
&lt;li&gt;The 23 mantissa bits stored in memory will look like this: 00100000000000000000000&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id=&quot;2-how-the-mantissa-maps-to-decimal-values&quot;&gt;2. How the Mantissa maps to Decimal Values&lt;/h3&gt;
&lt;p&gt;Each of the 23 bits represents a negative power of 2, starting right after the binary point:
$$\text{Value} = 1 + (b_1 \times 2^{-1}) + (b_2 \times 2^{-2}) + (b_3 \times 2^{-3}) + \dots + (b_{23} \times 2^{-23})$$
Where b₁ is the first mantissa bit, b₂ is the second, and so on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Bit 1 = 2⁻¹ = 0.5&lt;/li&gt;
&lt;li&gt;Bit 2 = 2⁻² = 0.25&lt;/li&gt;
&lt;li&gt;Bit 3 = 2⁻³ = 0.125&lt;/li&gt;
&lt;li&gt;Bit 23 = 2⁻²³ = 0.0000001192…&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is why a float has ~7 decimal digits of precision. The smallest step change the mantissa can make is 2⁻²³, which is roughly 1.19 × 10⁻⁷.&lt;/p&gt;
&lt;h3 id=&quot;3-the-exception-subnormal-numbers-stored-exponent-0&quot;&gt;3. The Exception: Subnormal Numbers (Stored Exponent = 0)&lt;/h3&gt;
&lt;p&gt;When a number gets so tiny that the exponent cannot go any lower, the float switches to subnormal mode to prevent dropping straight to zero.
When this happens:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The implicit leading bit changes from 1. to 0..&lt;/li&gt;
&lt;li&gt;The number is evaluated as: 0.xxxxxxxxxxxxxxxxxxxxxxx₂ × 2⁻¹²⁶.&lt;/li&gt;
&lt;li&gt;As the number gets smaller, more leading zeros creep into the 23 mantissa bits, causing you to steadily lose precision until no bits are left.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr /&gt;
&lt;h2 id=&quot;how-is-the-range-1-4-x-10-45-to-3-4-x-1038-of-float-computed&quot;&gt;How is the range &lt;code&gt;±1.4 × 10⁻⁴⁵ to ±3.4 × 10³⁸&lt;/code&gt; of &lt;code&gt;float&lt;/code&gt; computed?&lt;/h2&gt;
&lt;p&gt;The numbers $3.4 \times 10^{38}$ and $1.4 \times 10^{-45}$ are computed by converting the largest and smallest possible binary values of a float into base-10 (decimal) numbers.
To understand the math, we first look at how the 8 exponent bits work under the &lt;a rel=&quot;external&quot; href=&quot;https://en.wikipedia.org/wiki/IEEE_754&quot;&gt;IEEE 754 standard&lt;/a&gt;, and then factor in the 23 mantissa bits.&lt;/p&gt;
&lt;h3 id=&quot;step-1-the-8-bit-exponent-range-and-bias&quot;&gt;Step 1: The 8-Bit Exponent Range and Bias&lt;/h3&gt;
&lt;p&gt;An 8-bit binary number can represent integers from 0 to 255. However, IEEE 754 reserves two of these values for special purposes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;0 is reserved for zero and subnormal numbers.&lt;/li&gt;
&lt;li&gt;255 is reserved for Infinity and NaN (Not a Number).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This leaves a usable range of 1 to 254. To allow for negative exponents (tiny fractions), a bias of 127 is subtracted from the stored exponent.
$$\text{True Exponent} = \text{Stored Exponent} - 127$$&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Maximum Normal Exponent: $254 - 127 =$ $+127$&lt;/li&gt;
&lt;li&gt;Minimum Normal Exponent: $1 - 127 =$ $-126$&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;step-2-computing-the-max-range-3-4-times-10-38&quot;&gt;Step 2: Computing the Max Range ($3.4 \times 10^{38}$)&lt;/h3&gt;
&lt;p&gt;To get the absolute largest number, we maximize both the exponent and the mantissa:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Max Exponent: $2^{127}$&lt;/li&gt;
&lt;li&gt;Max Mantissa: All 23 bits are set to 1. In binary scientific notation, this represents $1.1111…_2$, which is effectively just a fraction below $2$ (specifically, $2 - 2^{-23} \approx 1.99999988$).&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Now, we multiply them together:
$$\text{Max Value} \approx 2 \times 2^{127} = 2^{128}$$
Using logarithms to convert $2^{128}$ to base-10:
$$2^{128} = 10^{128 \times \log_{10}(2)} \approx 10^{128 \times 0.30103} \approx 10^{38.53}$$
The exact math yields $3.4028235 \times 10^{38}$.&lt;/p&gt;
&lt;h3 id=&quot;step-3-computing-the-min-range-1-4-times-10-45&quot;&gt;Step 3: Computing the Min Range ($1.4 \times 10^{-45}$)&lt;/h3&gt;
&lt;p&gt;The value $1.4 \times 10^{-45}$ represents the absolute smallest non-zero positive number a float can hold. To achieve this, the system shifts from a “normal” number to a subnormal number:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Min Exponent: When the stored exponent is 0, the true exponent is fixed at $-126$.&lt;/li&gt;
&lt;li&gt;Min Mantissa: Instead of the usual implicit leading 1, the implicit leading digit becomes 0. We then set only the very last bit of the 23-bit mantissa to 1. This gives a value of $2^{-23}$.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Now, we multiply them together:
$$\text{Min Value} = 2^{-23} \times 2^{-126} = 2^{-149}$$
Using logarithms to convert $2^{-149}$ to base-10:
$$2^{-149} = 10^{-149 \times \log_{10}(2)} \approx 10^{-149 \times 0.30103} \approx 10^{-44.85}$$
The exact math yields $1.4012985 \times 10^{-45}$.&lt;/p&gt;
&lt;hr /&gt;
&lt;h2 id=&quot;the-classic-0-1-0-2-0-3-problem&quot;&gt;The classic &lt;code&gt;0.1 + 0.2 != 0.3&lt;/code&gt; problem&lt;/h2&gt;
&lt;p&gt;The reason 0.1 + 0.2 equals 0.30000000000000004 comes down to a fundamental limitation of computers: they think in base-2 (binary), while humans think in base-10 (decimal).
Because computers have a fixed size (usually a 64-bit double in languages like JavaScript, Python, and Java), they cannot accurately represent fractions that repeat infinitely in binary.
Here is the step-by-step breakdown of how this math error happens under the hood.&lt;/p&gt;
&lt;h3 id=&quot;1-the-repeating-fraction-problem&quot;&gt;1. The Repeating Fraction Problem&lt;/h3&gt;
&lt;p&gt;In base-10, you cannot cleanly write $1/3$ as a decimal. It becomes an infinite repeating fraction: 0.333333…. No matter how many 3s you write down on a piece of paper, you will eventually run out of space and be forced to round it.&lt;/p&gt;
&lt;p&gt;Binary has the exact same issue, but with different numbers. A fraction can only be written cleanly in binary if its denominator’s prime factors are powers of 2.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Clean in decimal: $1/10$ (0.1) and $2/10$ (0.2)&lt;/li&gt;
&lt;li&gt;Infinite in binary: Both 0.1 and 0.2 become infinite repeating binary fractions!&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;$$\text{0.1 in binary} = 0.00011001100110011…_2 \text{ (repeating forever)}$$
$$\text{0.2 in binary} = 0.00110011001100110…_2 \text{ (repeating forever)}$$&lt;/p&gt;
&lt;h3 id=&quot;2-computers-cut-off-and-round&quot;&gt;2. Computers Cut Off and Round&lt;/h3&gt;
&lt;p&gt;Since a 64-bit double only allocates 52 bits for the mantissa (precision), the computer has to chop off the infinite tail of these numbers and round them to the nearest possible binary value.
When you type 0.1 and 0.2, the computer actually stores slightly inaccurate values:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Stored 0.1 $\approx$ 0.10000000000000000555111512312578…&lt;/li&gt;
&lt;li&gt;Stored 0.2 $\approx$ 0.20000000000000001110223024625156…&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;3-the-math-and-the-4-at-the-end&quot;&gt;3. The Math and the “4” at the End&lt;/h3&gt;
&lt;p&gt;When the CPU adds these two rounded binary numbers together, the small rounding errors stack on top of each other.
If we look at the exact underlying decimal math the computer is doing:&lt;/p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#000000, #E6EDF3); background-color: light-dark(#FFFFFF, #0D1117);&quot; &gt;&lt;code data-lang=&quot;plain&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;  0.100000000000000005551115... (Stored 0.1)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;+ 0.200000000000000011102230... (Stored 0.2)&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;---------------------------------&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;  0.300000000000000016653345... (Exact sum)&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The exact sum ends in …16653345…. When the computer tries to fit this result back into a 64-bit double, it rounds it to the closest valid floating-point number it can possibly represent.&lt;/p&gt;
&lt;p&gt;That closest number happens to be exactly 0.30000000000000004440892098500626….&lt;/p&gt;
&lt;p&gt;When your programming language prints it out, it truncates the view to 17 digits, leaving you with the famous 0.30000000000000004.&lt;/p&gt;
&lt;h3 id=&quot;how-to-fix-this-in-code&quot;&gt;How to Fix This in Code&lt;/h3&gt;
&lt;p&gt;If you are dealing with everyday rounding or money, you don’t want these errors messing up your software. Here are the common solutions:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Format/Round for display&lt;/strong&gt;: Keep the math as-is, but round the string output when showing it to users (e.g., result.toFixed(2) in JavaScript).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Work in Cents (Integers)&lt;/strong&gt;: Instead of using decimals for money, multiply by 100 and work strictly with whole numbers (integers), since 10 + 20 = 30 is always exact.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Use Arbitrary-Precision Libraries&lt;/strong&gt;: Use built-in types designed for exact decimal math, such as BigDecimal in Java, decimal in Python or C#, or libraries like big.js in JavaScript.&lt;/li&gt;
&lt;/ol&gt;
</content>
        
    </entry>
</feed>
