Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Bits, Bytes, and Numbers

Open in Colab

In order to do math on a computer, you should have some idea of how computers represent numbers and do computations.

Unlike strongly typed languages (e.g. C), you don’t have to worry too much about type if you’re just scripting. However, you will have to think about this for many algorithms in scientific computing.

Bits and Bytes

A bit is a 0/1 value, and a byte is 8 bits. Most modern computers are 64-bit architectures on which Python 3 will use 64-bits to represent numbers. Some computers may be 32-bit architectures, and Python may use 32-bits to represent numbers - beware!

You can represent strings of bits using the 0b prefix. Be default, these will be interpreted as integers written in base-2. For example, on a 32-bit system,

0b101 = 00000000000000000000000000000101 (base 2) = 5 (base 10)

It is often easier to deal with hexadecimal (base-16), denoted with the 0x prefix

Let’s count in hexadecimal and binary:

By using the 0o prefix, you can also use octal (base 8) if you’d like

hexadecimal is often used because it breaks up bits into blocks of 4 (16 = 2^4). So a 64-bit type has some representation as a length-16 string in hexadecimal.

Integers

Integers are represented in base-2 using bits. Most modern computers are 64-bit architectures, and Python 3 will by default use 64-bit integers. However, unlike some languages, Python will use arbitrary precision so you don’t run into overflow errors:

However, when we call code written in C/C++ or fortran such as numpy, you can run into overflow issues

Bitwise operations

You can perform operations on bit strings in Python. To start with, we recall operations in boolean algebra: & (and), | (or), and ^ (xor) ~ (not).

Here’s a list of possible values for &

Performing an operation bitwise just performs the operation on each set of bit grouped by position (i.e. you can read off each column of the below using the table above)

  0b1100
& 0b0101
--------
  0b0100

Note that & and and are not equivalent. and does not operate bitwise - it interprets both as logical values (where 0 is False and any other number is True).

Two’s Complement

You’ll notice in the above, that ~ has a potentially unexpected output. We might have expected something like the following:

To explain why, we need to understand how negative integers are represented. Let’s consider a signed 8-bit integer. The first bit is the sign bit (0 indicates the number is positive, and 1 indicates the number is negative).

What about a negative number?

Naively, if we ignore the sign bit, -1 looks like "-"127, and -127 looks like "-"1. This is because negative integers are using two’s complement to represent negative integers, so you can’t just read off the number by ignoring the sign bit in the way you might expect.

You can compute the two’s complement of a number by inverting bits using a bit-wise not operation and adding 1 (ignoring integer overflow).

The two’s complement operation is its own inverse.

Note that positive numbers can go up to 127, but negative numbers go down to -128.

Why use two’s complement? The reason why is that using this representation you can use the same circuits in hardware for addition, subtraction, and multiplication of negative integers that you can for positive integers.

Floating Point Numbers

Real numbers are typically represented as floating point numbers on a computer. Almost all real numbers must be approximated, which means you can’t always ask for exact equality

The approximation error is called machine precision, typically denoted ϵ\epsilon

32-bit floating point numbers corrsepond to a float in C, and are also known as single precision numbers. 64-bit floating point numbers correspond to a double in C, and are also known as double precision numbers. 16-bit floats are half-precision, and 128-bit floats are quad-precision.

Double precision is the standard for many numerical codes. Quad- (or higher) precision is sometimes useful. A big trend in deep learning is to use lower-precision formats.

Floating point numbers are numbers written in scientific notation (in base-2). They contain a sign bit, a set of bits for the exponent, and a set of bits for the decimal (called the significand or mantissa).

For example, float32 has 1 bit for the sign, 8 bits for the exponent, and 23 bits for the mantissa.

img

For further reading on potential considerations with floating point numbers, see the Python documentation.

You can inspect the bits used in a floating point number in python using the bitstring package

(pycourse) $ conda install bitstring -c conda-forge

(this notebook installs it with pip if it is missing, e.g. on Google Colab)

Converting From Floating Point

If you want to convert the binary floating point representation to a decimal, you can read off the sign, exponent, and significand independently.

The sign bit is straightforward (0 is positive, 1 is negative). For floats, there are two equivalent zeros.

The exponent has a “bias” equal to half the number of possible bits. So to get the value of an 8-bit exponent, we subtract 2**7-1 = 127 = 0b01111111. So the exponent part is

2(b31b30…b23)2−1272^{(b_{31}b_{30}\dots b_{23})_2 - 127}

The decimal multiplies the exponential by

1+∑ib23−i2−i1 + \sum_{i} b_{23-i} 2^{-i}

Caution!

You can’t count on floating point numbers to count:

You must also be aware of accumulating rounding errors. Let’s test that by taking an average of many duplicates of 1/3.

There other ways to compute means, but they don’t do much better

You obtain the number by raising the significand to the base multiplied by the exponent.

It is often convenient to format floating point numbers for printing without showing full precision. An explanation of available options can be found in the format specification mini-language documentation. We’ll cover a few examples.

When formatting a floating point number with format, you put format specification in the curly braces {}, as in "{:width.precision}}.

The width denotes the total field width.

The precision denotes how many digits should be displayed after the decimal.

Exercise

  1. What is log⁡2(ϵ)\log_2(\epsilon) for 32 and 64-bit floats? (hint: use np.log2)

  2. How many bits do you think are used to represent the decimal part of a number in both these formats?

  3. If you take into account a sign bit, how many bits are left for the exponent?

  4. What is the largest exponent you can have for 32- and 64-bit floats? Keep in mind that you should have an equal number of positive and negative expoenents.

  5. Design an experiment to check your answer to part 4.

Check your answer with e.g. np.finfo(np.float32).max)

Notebook Cell
Notebook Cell
Notebook Cell
Notebook Cell
Notebook Cell
Notebook Cell