resource sheet
How I would learn Python as an MD/PhD student
if I had to do it again
Someone asked me this in the DMs, so here is the long version: what I did, in the order I did it, including the parts that were slower than they needed to be.
But should you actually? Sometimes, for a quick analysis, you will get away with an Excel plot. But if the data gets a bit more complex, or you end up with 2.5 billion data points, Excel won't even open that file, and SPSS and Prism are probably just going to crash. So there is no way around code-based data analysis if you work in clinical science or genomics.
But there isn't only Python (the language I focused on). If you will focus on statistics for clinical studies, R might be better for you. If you are interested in a broader approach, I would stick with Python.
1Where I am
I am a medical student, not a computer science student. I do not try to hold syntax in my head; this is not my specialty, so I mostly delegate it to an LLM of my choice. What I train instead is three things.
Structuring data in databases rather than in many loose files, in order to allow for good (machine) readability. You might want to look into SQL (Structured Query Language) for more complex data structures as well.
Reading code, in order to understand the generated code: which data is pulled from my database and how it is computed, what functions were used and how they are modified.
And writing pseudocode: describing exactly what I want, in what order, and why, before any of it is real Python. On the way of execution and optimisation I happily take feedback from AI. On the maths or statistics I want to stay on top.
Turning that description into working and documented syntax is the part I hand over. An LLM does it faster and cleaner than I ever will, and that gap is not closing in my favour. What stays mine is knowing what to ask for and being able to check the answer.
2My path, in order
3The most important courses I did
All openHPI, the open course platform of the Hasso Plattner Institute, a renowned German computer science university centre. Free, self-paced, no payment anywhere in the flow.
- Python for Beginners: English. The one to start with.
- Python: schnell und intensiv Programmieren lernen: German, if that is easier for you.
- Data Science Bootcamp: English. Where Python starts being useful rather than just correct.
4Like any new language, you need to stay consistent
Codewars is the Duolingo approach. Small problems, instant feedback, a streak to protect, however a bit difficult in the beginning. It is the easiest way to build a daily habit, but it has no specific direction. You can always ask an LLM for help to break down your problem. Here you can learn to structure a project. But try on your own first.
If you would rather have one long structured run, Harvard's CS50P is free on YouTube, about thirteen and a half hours, and will go through a lot.
5The step that helped me the most
Take a dataset. One of your own if you have one, otherwise the Titanic dataset, because it is everywhere and every tutorial uses it, so you will always find help when you get stuck. Then, in this order:
6Then generate the takeaway and make it visible
During my internship at BCG, our daily bread was to generate insights and derive strategies from them. But no matter how good (or bad) my model was, I always had to explain my calculations to people who didn't really care about the model. So the data needed to be interesting, visually appealing, and show the right data to underline your conclusions.
Always check if the plot is right for your purpose. I still see too many dynamite plots for biological tests. For a Q1 revenue report they show exactly the number, but in biology they hide the variance of your 10 mice.
For mice I always use scatterplots, either with the mean or a box plot. And I pretty much always run two-sided Welch tests. But this fits my sample size and endpoints. For clinical cohorts the pure scatterplot might not be misleading, but it is not insightful enough.
Also, this is the time to choose your colours for the thesis. Set a palette of 12-16 hex codes and embed them deeply in your project.
As a reference, the top is a typical consulting slide, focussing on 6 bar charts but doing so with absolute numbers and no variance. So this type is okay, however it misses the statistical tests, which is okay for industry but unacceptable for science. The middle gives more comparisons but doesn't get to the point, and it is also the layouting mess I see way too often in basic science. The third one aims to bridge the gap between those two.
Ace your presentations
In my first year of my PhD I tried to plot everything
Now I focus on 4 plots and the rest goes into the supplement. And my titles for presentations are always a question or an action title.
matplotlib when you want control over every element. seaborn when you want to write a statistical plot in one line. plotly when you want it to be interactive.
Plot the same result multiple ways (you won't one-shot it), then decide which one is honest. That decision is a skill of its own, and it is not a technical one.
7Where AI actually fits
It will make the process faster for you. Which is exactly why the skill worth having is specification: which library, which model, which plot, which axis, which test. A vague prompt gets you a plausible chart that answers a question you never asked, or forgets the important things.
So read the documentation of the libraries you use. It is free, it is the best reading material available, and it is what makes your prompts precise enough to be worth trusting.
And do not forget to read the code behind the plot. In my early tries, ChatGPT pulled the wrong dataset, or a variable was missing. That would be horrible in a paper. Also, do not just use the chat window: use Claude Code or Codex in VS Code, or Gemini in Google Colab.
8Not yet: certificates
Everything in this sheet is free or free to audit. A certificate is the easiest thing in this field to buy and the least interesting thing to show. One small project you can defend in conversation beats a row of badges.
every link, in order
The list
Python basics
| What | Why | Cost | Link |
|---|---|---|---|
| openHPI Python for Beginners | English, self-paced. The one I started with. | free | open.hpi.de/learn/python2025 |
| openHPI Python schnell & intensiv | The same idea in German. | free | open.hpi.de/learn/python2024 |
| Codewars | Daily reps. Gamified, easy to keep up, slow. | free | codewars.com |
| CS50P | Harvard, about 13.5 h, one long structured run. | free | youtube.com/watch?v=nLRL_NcnK-4 |
Then data science
| What | Why | Cost | Link |
|---|---|---|---|
| openHPI Data Science Bootcamp | Where Python starts being useful. | free | open.hpi.de/learn/datascience2023 |
| Titanic dataset | The practice dataset everyone uses, so help is everywhere. Also built into seaborn: sns.load_dataset("titanic"). | free | kaggle.com/c/titanic |
| NCATS OpenData | Public data if you want something less rehearsed. | free | opendata.ncats.nih.gov/rwd |
| IBM Databases & SQL | SQL for data science with Python. Certificate is paid, material is not. | audit free | coursera.org/learn/sql-data-science |
The libraries you probably will not get around
| Library | Use it for | Cost | Link |
|---|---|---|---|
| pandas | Loading and reshaping the data. | free | pandas.pydata.org/docs |
| statsmodels and scipy | The statistical models and the tests. | free | statsmodels.org · docs.scipy.org |
| scikit-learn | Classical machine learning. The fair comparison. | free | scikit-learn.org |
| TensorFlow | The deep-learning version of the same question. | free | tensorflow.org |
| matplotlib seaborn, plotly | Control, speed, interactivity, in that order. | free | matplotlib.org · seaborn.pydata.org · plotly.com/python |
Later, if you want to go further
| What | Why | Cost | Link |
|---|---|---|---|
| openHPI Time Series Analysis | One of the most common and most annoying data types. | free | open.hpi.de/learn/timeseries2025 |
| openHPI Efficient AI for Weather Forecasting |
Applied machine learning on a real physical problem. | free | open.hpi.de/learn/weatherpredictions2025 |
| Google Data Analytics | Broad, structured, aimed at jobs. Certificate is paid. | audit free | coursera.org/google-certificates/data-analytics-certificate |
| AWS Skill Builder | If the cloud side becomes relevant. | free tier | skillbuilder.aws |
Not from my path, but worth knowing
| What | Why | Cost | Link |
|---|---|---|---|
| Kaggle Learn | Short practical micro-courses on Python, pandas and visualisation. | free | kaggle.com/learn |
| Exercism Python track | 146 exercises, 17 concepts, automatic analysis and volunteer mentoring. | free | exercism.org/tracks/python |
| The official Python tutorial | Dry, short, and the definitive answer to how something actually works. | free | docs.python.org/3/tutorial |
newsletter
Want to stay updated?
New guides and resources go to my newsletter first. Sign up and you will get them as soon as they are out.