felix et al.

resource sheet

How I would learn Python as an MD/PhD student
if I had to do it again

the question in the DMs

Someone asked me this in the DMs, so here is the long version: what I did, in the order I did it, including the parts that were slower than they needed to be.

But should you actually? Sometimes, for a quick analysis, you will get away with an Excel plot. But if the data gets a bit more complex, or you end up with 2.5 billion data points, Excel won't even open that file, and SPSS and Prism are probably just going to crash. So there is no way around code-based data analysis if you work in clinical science or genomics.

But there isn't only Python (the language I focused on). If you will focus on statistics for clinical studies, R might be better for you. If you are interested in a broader approach, I would stick with Python.

1Where I am

I am a medical student, not a computer science student. I do not try to hold syntax in my head; this is not my specialty, so I mostly delegate it to an LLM of my choice. What I train instead is three things.

Structuring data in databases rather than in many loose files, in order to allow for good (machine) readability. You might want to look into SQL (Structured Query Language) for more complex data structures as well.

Reading code, in order to understand the generated code: which data is pulled from my database and how it is computed, what functions were used and how they are modified.

And writing pseudocode: describing exactly what I want, in what order, and why, before any of it is real Python. On the way of execution and optimisation I happily take feedback from AI. On the maths or statistics I want to stay on top.

Turning that description into working and documented syntax is the part I hand over. An LLM does it faster and cleaner than I ever will, and that gap is not closing in my favour. What stays mine is knowing what to ask for and being able to check the answer.

This is not an invite to skip the basics. Without syntax you cannot tell a working script from a confidently wrong one, and the confidently wrong one looks exactly the same.

2My path, in order

1Basics of syntax and how functions workopenHPI Python course, then Data Science, a lot of YouTube, discussions with ChatGPT to break down existing code I did not understand.
2A real dataset instead of exercisesTitanic, or one of my own.
3Statistical analysis of that datasetscipy, statsmodels.
4Visualising the result several waysmatplotlib, seaborn.
5The same dataset, but machine learningscikit-learn, TensorFlow.
6Comparing the two honestlyThe learning step.
7Reading library documentation properlyI did that too late and it is still underrated. It gives you ideas on how to use libraries and gives you ideas on new approaches. It also shows limitations and workarounds.
8A lot of trying around afterwardsI maybe should have done more courses, but I always had enough projects to work on and to try to learn from.

3The most important courses I did

All openHPI, the open course platform of the Hasso Plattner Institute, a renowned German computer science university centre. Free, self-paced, no payment anywhere in the flow.

  • Python for Beginners: English. The one to start with.
  • Python: schnell und intensiv Programmieren lernen: German, if that is easier for you.
  • Data Science Bootcamp: English. Where Python starts being useful rather than just correct.
openHPI is academically solid and sometimes fairly theoretical. That is only a strength if you want to understand what is happening underneath, and a risk if you just want to get going. Do the basics, then leave and build something. You will not get fluent by watching, but the courses contain a lot of exercises.

4Like any new language, you need to stay consistent

Codewars is the Duolingo approach. Small problems, instant feedback, a streak to protect, however a bit difficult in the beginning. It is the easiest way to build a daily habit, but it has no specific direction. You can always ask an LLM for help to break down your problem. Here you can learn to structure a project. But try on your own first.

If you would rather have one long structured run, Harvard's CS50P is free on YouTube, about thirteen and a half hours, and will go through a lot.

5The step that helped me the most

Take a dataset. One of your own if you have one, otherwise the Titanic dataset, because it is everywhere and every tutorial uses it, so you will always find help when you get stuck. Then, in this order:

a
Analyse and visualise the data before testing any hypothesis on it. Find the variables and the missing data points, and run descriptive statistics first. What does the data actually tell me?
b
What are the questions you could ask, what is the right test, what are the necessary assumptions, and are they met? Check what statistical test is appropriate for your data, or if it needs to be transformed beforehand. I started with scipy and statsmodels. If you study anything quantitative, this is your time to shine. Run regressions or correlation maps. I used this dataset as a playground.
c
Then throw a machine learning model at the same data. scikit-learn first, TensorFlow if you want the deep-learning version.
d
Compare them. Where do they agree? Where does the ML model just look more impressive without being better? That comparison is the moment statistics stops being an exam topic and starts being a tool.

6Then generate the takeaway and make it visible

During my internship at BCG, our daily bread was to generate insights and derive strategies from them. But no matter how good (or bad) my model was, I always had to explain my calculations to people who didn't really care about the model. So the data needed to be interesting, visually appealing, and show the right data to underline your conclusions.

Always check if the plot is right for your purpose. I still see too many dynamite plots for biological tests. For a Q1 revenue report they show exactly the number, but in biology they hide the variance of your 10 mice.

For mice I always use scatterplots, either with the mean or a box plot. And I pretty much always run two-sided Welch tests. But this fits my sample size and endpoints. For clinical cohorts the pure scatterplot might not be misleading, but it is not insightful enough.

Also, this is the time to choose your colours for the thesis. Set a palette of 12-16 hex codes and embed them deeply in your project.

As a reference, the top is a typical consulting slide, focussing on 6 bar charts but doing so with absolute numbers and no variance. So this type is okay, however it misses the statistical tests, which is okay for industry but unacceptable for science. The middle gives more comparisons but doesn't get to the point, and it is also the layouting mess I see way too often in basic science. The third one aims to bridge the gap between those two.

Disclaimer: all three slides were made with Claude, with rather vague prompts, for visualisation only. Sadly, right now I have to focus too much on my 2nd state exam and my PhD to make that illustration by hand.

Ace your presentations

consulting slide
Consulting. Clear, but no test.
typical basic-science slide
Basic science. Everything reported, nothing concluded.
bridge slide
The bridge. Same numbers, tested and interpreted.
Always explain how you used AI, and write it in any publication or scientific paper. It is allowed to use AI to support your tasks in most cases. What is not allowed is to let AI take over the thinking. (Check the last page for the full disclaimer of this document.)

In my first year of my PhD I tried to plot everything

Now I focus on 4 plots and the rest goes into the supplement. And my titles for presentations are always a question or an action title.

matplotlib when you want control over every element. seaborn when you want to write a statistical plot in one line. plotly when you want it to be interactive.

Plot the same result multiple ways (you won't one-shot it), then decide which one is honest. That decision is a skill of its own, and it is not a technical one.

7Where AI actually fits

It will make the process faster for you. Which is exactly why the skill worth having is specification: which library, which model, which plot, which axis, which test. A vague prompt gets you a plausible chart that answers a question you never asked, or forgets the important things.

So read the documentation of the libraries you use. It is free, it is the best reading material available, and it is what makes your prompts precise enough to be worth trusting.

And do not forget to read the code behind the plot. In my early tries, ChatGPT pulled the wrong dataset, or a variable was missing. That would be horrible in a paper. Also, do not just use the chat window: use Claude Code or Codex in VS Code, or Gemini in Google Colab.

Always check how your data is shared with AI, and which read and write permissions it has. Some data is confidential. The code might be allowed to be generated with AI, but the output, and especially the raw data, may never leave the system. Enterprise options may offer better data security, but you should always check your university's or hospital's policies. I am at a German and an American institution, so depending on the project GDPR (DSGVO) or HIPAA applies. Those are still the most important rules.

8Not yet: certificates

Everything in this sheet is free or free to audit. A certificate is the easiest thing in this field to buy and the least interesting thing to show. One small project you can defend in conversation beats a row of badges.

every link, in order

The list

Python basics

WhatWhyCostLink
openHPI
Python for Beginners
English, self-paced. The one I started with.free open.hpi.de/learn/python2025
openHPI
Python schnell & intensiv
The same idea in German.free open.hpi.de/learn/python2024
CodewarsDaily reps. Gamified, easy to keep up, slow.free codewars.com
CS50PHarvard, about 13.5 h, one long structured run.free youtube.com/watch?v=nLRL_NcnK-4

Then data science

WhatWhyCostLink
openHPI
Data Science Bootcamp
Where Python starts being useful.free open.hpi.de/learn/datascience2023
Titanic datasetThe practice dataset everyone uses, so help is everywhere. Also built into seaborn: sns.load_dataset("titanic").free kaggle.com/c/titanic
NCATS OpenDataPublic data if you want something less rehearsed.free opendata.ncats.nih.gov/rwd
IBM
Databases & SQL
SQL for data science with Python. Certificate is paid, material is not. coursera.org/learn/sql-data-science

The libraries you probably will not get around

LibraryUse it forCostLink
pandasLoading and reshaping the data. free pandas.pydata.org/docs
statsmodels
and scipy
The statistical models and the tests.free statsmodels.org · docs.scipy.org
scikit-learnClassical machine learning. The fair comparison.free scikit-learn.org
TensorFlowThe deep-learning version of the same question.free tensorflow.org
matplotlib
seaborn, plotly
Control, speed, interactivity, in that order.free matplotlib.org · seaborn.pydata.org · plotly.com/python

Later, if you want to go further

WhatWhyCostLink
openHPI
Time Series Analysis
One of the most common and most annoying data types.free open.hpi.de/learn/timeseries2025
openHPI
Efficient AI for Weather Forecasting
Applied machine learning on a real physical problem. free open.hpi.de/learn/weatherpredictions2025
Google
Data Analytics
Broad, structured, aimed at jobs. Certificate is paid. coursera.org/google-certificates/data-analytics-certificate
AWS Skill BuilderIf the cloud side becomes relevant.free tier skillbuilder.aws

Not from my path, but worth knowing

WhatWhyCostLink
Kaggle LearnShort practical micro-courses on Python, pandas and visualisation.free kaggle.com/learn
Exercism
Python track
146 exercises, 17 concepts, automatic analysis and volunteer mentoring.free exercism.org/tracks/python
The official
Python tutorial
Dry, short, and the definitive answer to how something actually works.free docs.python.org/3/tutorial

newsletter

Want to stay updated?

New guides and resources go to my newsletter first. Sign up and you will get them as soon as they are out.

Sign up for the newsletter

Course details checked 19 August 2026. openHPI courses run self-paced and are free to enrol. Coursera courses are free to audit; the certificate is not. Nothing here is sponsored.

Content and opinions are mine. Claude was used to: correct grammar and spelling; structure the text; verify the course details and check every link; lay out this document; and produce the three example slides in section 6, including the statistical analysis shown on them.