Guide
How I navigate AI in research as a student
A personal workflow for using AI while retaining responsibility for the research.
I use AI frequently in my research. I use it to write and debug code, structure dictated thoughts, organize information, explain unfamiliar concepts, and conduct preliminary explorations of data.
I have also worked with statistics for several years, but applied experience is different from formal statistical expertise. I try to recognize that boundary. When an analysis exceeds my training, I consult a statistician. I am also aware that in some cases a common AI tool may produce a better statistical analysis than I would. That is why I use it to consult, but still supervise. I also disclose relevant uses of AI to my supervisors, collaborators, and, where appropriate, in the manuscript.
To my surprise, more people than I expected believe that once the research question and data are available, an AI tool can determine and perform the statistical analysis. That approach concerns me. A PICO and an Excel file do not establish whether observations are paired, whether clustering matters, how missing data should be handled, or whether a model’s assumptions are satisfied.
If I cannot assess those decisions myself, a plausible-looking result may be difficult for me to challenge. This guide describes how I try to use AI without losing that critical step.
What I use AI for
I use AI to:
- structure thoughts that I have dictated;
- turn unstructured information into structured data;
- automate repetitive tasks;
- write and debug Python or R code;
- explain unfamiliar terminology;
- explore data before planning a formal analysis;
- prepare questions and possible next steps for discussion.
For voice to text: I always run into the problem that I type too slowly for my thoughts. By the time I finish typing a sentence, my head is already three sentences ahead, and that breaks the structure. For me, that is the most pivotal improvement. The second one is writing code, or automating, writing simple Python scripts for redundant tasks.
These are the tools I use. I have the luxury of having access to enterprise solutions for my AI tools.
ChatGPT
Especially for voice to text, and for my writing style and redacting my writing.
Claude
The same as ChatGPT: voice to text and writing.
Codex and Claude Code
For code. They are slightly different, but I like both of them.
Claude Science
For research questions and iterations. It connects Claude models to scientific databases, specialist tools and local computing environments, and keeps the code and provenance behind an analysis.
OpenEvidence
Not officially available in the EU, but in my opinion a great help in complementing my literature research.
Consensus
Works with plugins in both Claude and ChatGPT and helps to narrow down a literature search. Despite really liking the explanations of both, I would only use them as supplements.
Gemini and NotebookLM
For research that is more on the administrative side, processes. NotebookLM oftentimes has free student sign-ups.
Copilot
I barely use it, except to query my Outlook emails.
Wherever possible, I store data on my computer or an institutional server.
Extending my literature search
I use AI to explain terms, break down difficult concepts, and suggest analogies. I also use tools such as OpenEvidence to orient myself in medical literature and Consensus to identify potentially relevant publications.
These tools help me broaden a search and find useful starting points. I still open the papers, read the relevant sections, examine the methods and study population, and confirm that each source supports the claim for which I intend to cite it.
A clear explanation may still be incomplete. A real citation may still be irrelevant to the specific claim. I remain responsible for verifying both.
Patient and confidential data
I would not upload identifiable or confidential patient-level data to a general-purpose consumer chatbot.
Patients place their trust not only in individual researchers but also in the institutions responsible for protecting their information. Any use of sensitive data must therefore follow the relevant institutional policy, ethics approval, data-use agreement, access restrictions, and study-specific permissions.
For sensitive research, I would use AI only in an environment approved for the relevant data and purpose. When possible, I would work with synthetic data, a data dictionary, or summary output rather than raw patient-level data.
Local execution keeps the raw data on my computer. What I put into a prompt, and what the model sends back, is still uploaded and processed by the provider.
- Safest
Code only
The AI writes the code. I run it on the real data myself. To test it, the AI gets placeholder values: a similar distribution, a similar minimum and maximum, a similar number of missing values. The price is a lot of manual work to implement and fix everything.
AI sees: no real data
Local data, local agent
The data and Python stay on my machine, but an agent with file access can open the data, sometimes by accident.
AI sees: whatever it opens
Enterprise plan
The AI sees the data. Under the contract, it is not used for training.
AI sees: everything it is given
- Least safe
Copied into a chat
Full data access. Depending on the settings, it may be used for training. It is also not reproducible: I do not get to see the code that made the figure, and chatbots sometimes pull values from memory instead of from the dataset. I stay away from this.
AI sees: everything
Basic statistical decisions I need to understand
Before asking AI to implement an analysis, I try to specify:
- the endpoints and hypotheses;
- the comparison groups;
- whether observations are independent or paired;
- the statistical test or regression model;
- clustering, repeated measurements, and missing data;
- the assumptions that must be checked;
- any multiplicity or false discovery rate correction;
- whether testing is one- or two-sided;
- planned sensitivity analyses.
I fix my statistics before I start, not after the AI has run them and explained them to me. Re-evaluating afterwards impairs my statistical credibility.
AI can question my plan and identify possible omissions. I still need to understand and approve the methodological decisions that enter the paper.
1 Me
- specifies endpoints, groups and tests
- uses AI to discuss these on a broad, conceptual level
- fixes the statistics before anything runs
2 AI
- takes over the coding
- writes and debugs code
- structures dictated thoughts
- suggests starting points in the literature
- questions the plan and finds omissions
- proposes possible next steps
3 Me
- checks the sources behind my claims
- understands and approves the methods
- keeps the code and provenance
- discloses AI use
4 Statistician or supervisor
- takes on analyses beyond my training, or shows me how to do them and which tools to use
- advises on methodological decisions I am unsure about
- reviews previous and proposed steps
- checks the work
When I replicated the code for a figure, an AI tool at some point inserted a correction for multiple comparisons without consulting me. That is not necessarily wrong, but it altered the results. Its suitability depends on the hypotheses, the analysis plan, and the conclusions the study is intended to support. The concern is an unexplained statistical choice being incorporated because its output looks reasonable.
The same principle applies to one-sided testing. A one-sided test may be justified in specific settings, but the decision should be made and documented in advance rather than inferred after examining the results. When I am uncertain, I discuss the decision with a statistician. International guidance similarly expects multiplicity and sidedness to be addressed during analysis planning. ICH E9
I want the reported methods to be explainable and defensible by one of the human authors.
Reproducibility
An analysis is reproducible only if I retain the information needed to run it again. Depending on the workflow, this can include the code, data version, parameters, package versions, software environment, and methodological decisions.
I prefer working in RStudio, VS Code, Claude Science, or another environment where I can inspect the scripts and rerun each step. If the AI does not require the raw data, I can provide the data structure or selected outputs and execute its code locally.
Every final figure should remain connected to the code and source variables that produced it. If a value changes, I should be able to rerun the analysis rather than reconstruct the figure manually.
Using AI to prepare better questions
When I am uncertain about the next step, I use AI to identify possible approaches and help me formulate a proposal. I can then ask my supervisor something concrete:
I see two possible next steps. I currently prefer this one for these reasons, but I would like your feedback before proceeding.
In my experience, this leads to a more focused discussion. I have considered the options while still asking for expert review before making a consequential decision.
It means I no longer ask openly what I should do next, but proactively propose what to do. That shows the limited scope of my knowledge, and that is exactly the process where I learn. In my opinion, the job of a PhD student is not primarily to produce papers, but to learn. And the job of a PI is to teach me that and discuss it with me, and it is on me to ask for it.
Disclosure and responsibility
I disclose relevant AI use according to the requirements of the institution, collaborators, and target journal. If AI contributed to data analysis, code, or figure generation, I document its role with enough detail to support review and replication.
This is consistent with current ICMJE recommendations, which state that authors remain responsible for AI-assisted content and should disclose how AI was used. When AI contributes to conducting a study, its use should be described in the Methods section with sufficient detail for replication (ICMJE Recommendations, section IV.A.3.d).
It is important to disclose the use of tools, especially those that help in decision making. At the same time, I understand that some PIs are not as open to AI as others, and that this can make a good research environment difficult. I see that as another thing we should work on.
My standard is whether I can explain my methodological choices, reproduce the figures, trace claims to their sources, and defend the analysis to my collaborators and readers. I am ultimately responsible for the work to which I attach my name, and for the work others trust me to do properly.