Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Basic Containers and Packages

Open in Colab

Python Lists

Lists are a data structure provided as part of the Python language. Lists and more on lists..

A list is compound data type which is a mutable, indexed, and ordered collection of data.

Lists are often constructed using square brackets: [...]

Usually lists will be generated programatically. One way you can do this is by using the append method

or the extend method, which appends all elements in another list

You can also extend lists using the + operator

You can also generate lists using list comprehensions. Comprehensions are “Pythonic” which is a vauge term roughly meaning “something a Python programmer would write”.

Generally, comprensions consist of [expression loop conditional]

This looks a lot like set notation in mathematics. E.g. for the set

y={i∣i∈x,i≠4}y = \{i \mid i \in x, i \ne 4\}

we compute

Indexing

Python is 0-indexed (like C, unlike fortran/Matlab). This means a list of length n will have indices that start at 0, and end at n-1.

This is the reason why range(n) iterates through the range 0,...,n-1

you can access elements starting at the back of the array using negative integers. A good way to think of this is the index -1 translates to n-1

Slicing - you can use the colon character : to slice an array. The syntax is start:end:stride

Lists are mutable, which means you can change elements

Other Python Collections

There are other collections you might use in Python:

  • Tuples (...) are ordered, indexed, and immutable

  • Sets {...} are unordered, unindexed, and mutable

  • Dictionaries {...} are unordered, indexed, and mutable

These collections also support comprehensions.

You can find additional types of collections in the Collections module

Numpy

If you haven’t already:

conda install numpy

Numpy is perhaps the fundamental scientific computing package for Python - just about every other package for scientific computing uses it.

Numpy basically provides a ndarray type (n-dimensional array), and provides fast operations for arrays (i.e. compiled C or Fortran).

We’ll do some deeper dives into numpy in future lectures. For now, we’ll cover some basics. For those who want to dive in now, here are some tutorials

You can find lots of information in the numpy documentation

You can easily generate numpy arrays from list data

A 2-dimensional array can be generated by lists of lists

a few class members:

Other ways of obtaining numpy arrays:

Indexing

1-dimensional arrays are indexed in the same way as lists (0-indexed, can use slices, etc)

you can also index using lists of indices

2-dimensional arrays are a bit different from lists of lists:

you can also use slices, index sets, etc. in multi-dimensional arrays.

If you only provide 1 index, you’ll get the corresponding row (or set of rows if slicing)

Arithmetic

Numpy arrays support basic element-wise arithmetic, assuming arrays are the same shape.

Note: there are more complicated broadcasting rules for different-shaped arrays, which we’ll cover some other time.

Warning: The * operator applied to 2-dimensional arrays is not the same as matrix-matrix multiplication. It will perform element-wise multiplication instead.

Numpy provides the @ operator for matrix multiplication. You can also use the matmul() or dot() (dot product) methods.

Numpy provides a variety of mathematics functions that you can use with numpy arrays. Numpy is vectorized, meaning that it is typically much faster to perform array operations than to use explicit for loops. This should be a familiar concept to Matlab users.

PyPlot

PyPlot is a go-to plotting tool for Python. It is fully operable with numpy arrays.

conda install matplotlib

CSV files, Pandas

The *.csv extension is typically used to denote a “comma seperated value” file. These types of files are often used to store arrays in human-readable plain text.

Here’s an example:

0, 1, 2, 3
4, 5, 6, 7
...

You can save numpy arrays to files using np.savetxt()

Files can be loaded using np.loadtxt()

Often, scientific data has some meaning associated with numbers. In this case, the csv file might have a header, and every row is a different data point.

temperature, density, width, length
0, 1, 2, 3
4, 5, 6, 7
...

You can still load using numpy, but it is easy to loose track of what the different columns of the array mean.

The solution for this sort of data is to use a Pandas dataframe

conda install pandas

You can get columns of a dataframe by using the column label

To get rows, use the iloc parameter:

You can easily plot labeled columns