↓Skip to main content

All the Python You Need For Data Science, Part 1

·16 mins · Learning Python
Author
Alex

TLDR? Get the 99 project ideas and start building!

Python is important to know as a data scientist for a couple of reasons: it’s one of the main general purpose data science languages, and, it’s useful for scripting and automating. In this series of posts we’ll cover what Python you need to know to get started as a data scientist. This post covers the very basics, the foundation needed before we dive into using it for data science.

This post is very hands on, I recommend you open your favorite IDE and follow along.

Setting up your environment #

Python often comes installed on your machine, so you could just start writing and running code without any setup. This approach eventually leads to problems with Python versions and/or dependencies. I recommend that you learn how to set up a Python environment correctly once and for all, then script or template it, so that it’s easy to always set it up correctly. This will save you a lot of pain down the road.

Packages and projects #

I recommend using uv to manage Python projects and dependencies, because it’s fast and combines several different tools into one. (If you are feeling philosophical, consider what it means that it’s written in Rust.)

Once you have installed it, a very minimal uv workflow would be:

uv init mycoolproject && cd mycoolproject # set up a new project
uv run main.py # run the code, handling dependencies and environment

Behind the scenes, or more specifically, in configuration files in that directory, uv keeps track of what Python version and what dependencies you are using.

Running Python code #

To get the most out of this guide, it helps to follow along and type and run the code. The easiest way to do this is to use the REPL. The REPL (“read-eval-print loop”) is the interactive Python prompt: you type a line, Python runs it and prints the result. Start one inside your project with:

uv run python

The REPL is great for trying things out, but it doesn’t save what you type. Instead, you can write your code in a script and then load it into the REPL.

Here’s an example. Write a file like this, call it basics.py:

models = ["lr", "rf", "xgb"]
n_models = len(models)

and run it with the -i flag:

uv run python -i basics.py

This runs the script and then opens the REPL. All the variables from the script are defined, so you can explore them interactively:

>>> n_models
3
>>> models[0]
'lr'

When you change the script, you can reload it with:

>>> exec(open("basics.py").read())

This runs the file again in the current session, so your variables pick up the new code. This can be great for fast iteration. If you run into problems, just exit and restart the REPL using the command above. Most editors can also send the current line or selection to a running REPL. This is my favorite way of working interactively, but, beware that not all editors automatically use your project’s .venv. When you’re done with the REPL, type exit().

Code quality #

As you start writing Python, you can use the ruff tool to automatically check for common issues, and to format the code nicely. Doing this right from the start will help you maintain good habits when writing code.

uv add --dev ruff
uv run ruff check   # lint: unused imports, potential bugs
uv run ruff format  # format: one consistent style

Python basics #

In this section we will go through some basic Python concepts and building blocks. The first part focuses on key data structures, the second on assorted other concepts. In most data science applications, data is stored in specialized data structures provided by libraries (e.g. numpy, pandas). Nevertheless, data structures like lists and dictionaries are useful for handling everything else, like for example a list of models, a dictionary of experiments and their results. They are also helpful for learning and exploring algorithms (although once you need to implement algorithms by hand, you may be using more specialized data structures or other languages than Python).

We will use these structures through an example. We are going to simulate the revenue growth of a business. This business has a certain number of customers, and, gains a certain number of customers each month. It also loses a certain percentage of customers. Each customer contributes a certain amount of revenue.

Lists: simulate customer growth #

The simulation runs over time. We will store each time period (month) as a Python list. We loop over the months. Each month, the current number of customers decreases by 5%, and, increases by 100 customers. Each month gets appended as a new element in the list. The most important thing to notice about this type of loop is that it uses the state (current) from the previous iteration. This type of operation is sometimes called a scan or accumulate.

customers = []
current = 0
for month in range(12):
    current = current * 0.95 + 100
    customers.append(current)

Once we have the state in each month saved as elements of a list, we can loop over it and make various calculations. A simple example is just calculating revenue, which is the same for each customer for each time period. For this loop we use list comprehension. In this type of loop, we don’t have access to the state from the previous iteration: we are just applying the same calculation (* 20) to each element. This operation is sometimes called a map.

revenue = [c * 20 for c in customers]

When calculating month over month growth, we are comparing one month to the preceding month, so we do need to access more than one element at a time. In this example, we can achieve that by zipping together one list, with the same list just shifted one element. This is sometimes called a windowed map. Note that we can assign more than one variable at a time with the a, b in syntax. This is called “tuple unpacking” (a tuple is a fixed group of values, like (1, 2)).

growth = [b / a - 1 for a, b in zip(revenue, revenue[1:])]

Lists are most useful when you are mostly working through all elements in order, and, are less focused on specific elements.

Dictionaries: name the assumptions #

When you do need access to specific elements, dictionaries are a good alternative. They store key-value pairs, and, can on average look up a value given a key in O(1) time. Because they “map” keys to values, dictionaries are also sometimes, confusingly, called “maps”. Dictionaries are often implemented by taking the key and applying a hash function to it, which maps the key to a specific slot in memory.

We can extend our example by encoding the parameters of our simulation in a dictionary.

assumptions = {"new": 100, "churn": 0.05, "price": 20}

customers = []
current = 0
for month in range(12):
    current = current * (1 - assumptions["churn"]) + assumptions["new"]
    customers.append(current)

revenue = [c * assumptions["price"] for c in customers]

In this case we are only storing a few values, so there’s no speed gain. A good example of using dictionaries to get speed ups is database joins, where the smaller table’s values get put into a dictionary, that can then be looked up once for every value of the larger table.

Functions: wrap it up #

Functions wrap up related functionality, so that you can, for example, apply it many times, or to different inputs. They also just help organize code.

The return value of the function is indicated by the return statement. Without such a statement, the result of the function will be None no matter what it does. Defaults can be specified in the parameters list, like with months=12 in this case.

assumptions = {"new": 100, "churn": 0.05, "price": 20}


def simulate(new, churn, price, months=12):
    customers = []
    current = 0
    for _ in range(months):
        current = current * (1 - churn) + new
        customers.append(current)
    return [c * price for c in customers]

Sometimes you want to wrap up some functionality for a use case that is so specific that it does not need a name. You can create these throwaway functions with the lambda keyword. We will see an example when we sort scenarios below.

A shorthand for calling a function with many arguments is to use “argument unpacking”, indicated by **, which is when you call a function with a dictionary with the keys corresponding to the function’s named arguments.

revenue = simulate(**assumptions)

Wrapping up functionality into functions with appropriate arguments is a good way to be able to run your data analysis or modeling on different inputs. Keep in mind that variables defined inside a function are only in scope in that function. But, variables that are defined at a higher scope outside the function will still be available for reading inside the function. Assigning to it will create a new local variable.

Compare scenarios: a dictionary of dictionaries #

Now that we have a function for simulating revenue, we want to run many different scenarios and compare them.

We can do this by defining our scenarios as a nested dictionary, where the outer key is the scenario name, and then each value is a dictionary itself with the parameter values for that scenario.

# simulate() from the previous step

scenarios = {
    "base": {"new": 100, "churn": 0.05, "price": 20},
    "sticky": {"new": 100, "churn": 0.02, "price": 20},
    "bold": {"new": 150, "churn": 0.08, "price": 20},
}

We can run these scenarios with a dictionary comprehension, where we iterate over the scenarios dictionary, building a new one where the keys are the scenario names and the values are the results of the simulation.

Notice that items() returns a key-value pair, which we unpack into name and a.

results = {name: simulate(**a) for name, a in scenarios.items()}

We can then sort the result names, using the key syntax, where we use an anonymous function, defined with lambda, that returns the last month’s revenue (using [-1] to select the last element in the 12-month list of revenues) for each scenario.

ranked = sorted(results, key=lambda name: results[name][-1], reverse=True)
for r in ranked:
    print(r)

Records: a dataclass instead of a dict #

Dictionaries are great for access to specific elements, but, quickly become unwieldy. A better data structure is a record, which in Python is called a dataclass. A record is sometimes called a product type.

For the cost of a little more boilerplate, we get several advantages. First, everything is defined in one place. If you make a typo in a field name, that will be caught earlier (when you create the object). Defaults can be defined once, instead of having to be populated for every instance. Dataclasses can be frozen, so that they can’t be changed once created. This can help catch many bugs. The main drawback is that you need to know the field names when defining the record. Note the new: int syntax, where the colon and part following it is a “type annotation”, labeling the new field as holding an int. These annotations are just informative, Python does not enforce them.

Dataclasses are defined in the following way. Notice the use of the @dataclass line, which is a decorator, a function that takes a function or class and modifies it, in this case adding functionality to make class Scenario a dataclass. A simple example of a decorator would be a logging function, which takes a function and wraps it in a print or log statement.

from dataclasses import dataclass, replace


@dataclass(frozen=True)
class Scenario:
    new: int
    churn: float
    price: float = 20.0

Creating an instance of a dataclass with a field name that does not exist will raise an error. This way errors can be caught earlier.

Scenario(new=100, chrun=0.05)  # TypeError: unexpected keyword argument 'chrun'

Once we have defined Scenario, we can use it like this. Note that we are redefining the function simulate here, to take a Scenario instead of several separate parameters.

def simulate(s, months=12):
    customers = []
    current = 0.0
    for _ in range(months):
        current = current * (1 - s.churn) + s.new
        customers.append(current)
    return [c * s.price for c in customers]

base = Scenario(new=100, churn=0.05)

# create a new scenario by modifying an existing one.
# replace returns a new record.
sticky = replace(base, churn=0.02)

simulate(base)
simulate(sticky)

More building blocks #

In a post like this we are not aiming for a thorough, ground up walk through of all of Python. Rather, we are picking the parts that together get you to the first rung, where you can start using Python in your data science work. The goal is to get you started, so that you can be effective, and, hopefully whet your appetite to learn more about those parts that interest you.

So if this section has the feel of a grab-bag, it is because it is, a bag of useful tools.

Slicing and indexing #

While the canonical use case for a list is when you will be accessing all elements, Python provides lots of options for accessing specific elements, called indexing, and, accessing parts of the list, so called slicing. The syntax is lst[start:stop:step].

revenue = simulate(base)   # 12 monthly values

revenue[0]       # 2000.0   first month
revenue[-1]      # 18385.6  last month
revenue[:3]      # first quarter (stop is exclusive)
revenue[3:6]     # second quarter: start inclusive, stop exclusive
revenue[-3:]     # last three months
revenue[::3]     # every third month

quarters = [sum(revenue[i : i + 3]) for i in range(0, 12, 3)]
# [11605, 27065, 40320, 51684] (rounded)

Strings can also be sliced in the same way.

"revenue"[:3]    # 'rev'  (strings slice the same way)

"paris"[::-1] # reverse a string (or list) by stepping through it from the end.

Conditionals #

if, elif and else are used to test conditions and execute code depending on what condition is true. The first branch with a true condition will execute, the rest will be skipped. The else branch will execute if nothing above it has executed. Keep in mind that the code to be executed after a conditional has to be indented one level.

if revenue[-1] > revenue[0]:
    trend = "growing"
elif revenue[-1] < revenue[0]:
    trend = "shrinking"
else:
    trend = "flat"

trend = "growing" if revenue[-1] > revenue[0] else "not growing"   # simplified one-line version

Formatting strings #

Strings are enclosed by " or '. Prefixing with an f gives access to a formatting mini-language inside the quotes. Using {} you can execute any expression and insert the result into the string. Numbers can be formatted. To debug, use {x=} to show the value of any variable or expression.

# After the colon comes the format spec:
#   ,    thousands separator
#   .0f  fixed-point number with 0 decimals (.2f gives 2 decimals)
#   .1%  multiply by 100, 1 decimal, add a % sign
#   <8   left-align in a field 8 characters wide
#   >10  right-align in a field 10 characters wide
f"Month 12 revenue: {revenue[-1]:,.0f}"   # 'Month 12 revenue: 18,386'
f"Churn: {base.churn:.1%}"                # 'Churn: 5.0%'

# Debugging: a trailing = prints the expression itself, then its value
f"{base.new=}"                            # 'base.new=100'
f"{revenue[-1] / revenue[0]=:.1f}"        # 'revenue[-1] / revenue[0]=9.2' (format spec after =)

for name, s in {"base": base, "sticky": sticky}.items():
    print(f"{name:<8}{simulate(s)[-1]:>10,.0f}")
# base        18,386
# sticky      21,528

Errors and exceptions #

Once your code is running as expected, you might think that there is no need to handle any errors. Nevertheless, there are a couple of important situations when you might want to handle errors, or, even cause an error, which is often called “raise” or “throw” an error.

When do you want to handle errors that might come up in the code? In data science applications there are only a few situations when you want to handle errors and not cause the program to exit. In most cases, you want the program to crash on an error, because it indicates a problem with the data or model that you probably should fix.

One situation when it does help to handle errors is long-running loops over many different items, where a few bad ones shouldn’t kill everything. Examples of this are calling APIs that might error repeatedly, or, parsing many files.

As another example, say that simulate is an expensive function that can sometimes fail, and, you have a large number of scenarios that you want to test. Then you can call simulate inside a try ... except block. This will catch the error (stored in an Exception object) and save it to the failures dictionary rather than stopping the entire program.

raw_scenarios = [
    {"new": 100, "churn": 0.05},
    {"new": 100, "chrun": 0.05},      # typo in a field name
    {"new": 80, "churn": "high"},     # a string where a number should be
    {"new": 120, "churn": 0.08},
]

results = {}
failures = {}
for i, raw in enumerate(raw_scenarios):
    try:
        results[i] = simulate(Scenario(**raw))
    except Exception as e:  # one bad scenario shouldn't kill the whole batch
        failures[i] = f"{type(e).__name__}: {e}"

print(f"{len(results)} ok, {len(failures)} failed")   # 2 ok, 2 failed
for i, msg in failures.items():
    print(f"  scenario {i}: {msg}")
# scenario 1: TypeError: Scenario.__init__() got an unexpected keyword argument 'chrun'
# scenario 2: TypeError: unsupported operand type(s) for -: 'int' and 'str'

This approach saves you from having to rerun the entire list of scenarios again when trying to fix the bad ones. The other situation when you do want to handle errors is if you are serving a data science product, like a model, through an API. Then you absolutely have to have good error handling.

Except for these cases, we want to cause our program to fail with an error, so much that we sometimes are even ready to “raise” (create) an error. Why would we even want this? The idea is that we always want to “fail fast”, so that we can fix fast too. In data science our conclusions and outputs are only ever as good as our inputs. So we need to make sure the inputs are good, before we invest time and money into analysis or training a model. We can do this by “raising” errors, like in the example below, which will stop with a ValueError.

def validate(s):
    if not 0 <= s.churn <= 1:
        raise ValueError(f"churn must be between 0 and 1, got {s.churn}")
    return s

validate(Scenario(new=100, churn=5))

In this case, we make sure that the input churn is always between 0 and 1, and, stop the program if it is not. This prevents bad values from causing hard to diagnose (or even invisible) problems further downstream.

Reading and Writing Files #

Almost any non-trivial programming will involve reading and writing files. These are the basics of doing this with Python. We use pathlib to handle paths. Opening a file using with open(...) as f handles the opening and closing of the file for us.

from pathlib import Path

out = Path("output")
out.mkdir(exist_ok=True)

# Write plain text: one line per month
# "w" argument to open means open for writing
with open(out / "revenue.txt", "w") as f:
    for month, r in enumerate(revenue, start=1):
        f.write(f"{month},{r:.0f}\n")

# Read it back, one line at a time
with open(out / "revenue.txt") as f:
    for line in f:
        month, r = line.strip().split(",")
        print(month, r)                          # 1 2000, then 2 3900, ...

# For small files, read or write the whole thing in one go
text = (out / "revenue.txt").read_text()         # the whole file as one string
lines = text.splitlines()                        # a list of lines, without the "\n"
(out / "notes.txt").write_text("base scenario, 12 months\n")

Conclusion #

This guide provides a crash course in the very basic Python you need to start learning how to use for data science. That’s the topic of another post.