PhD Course on Large Language Models and AI Agents for Research Data Analysis
Schedule and location
Aalto University School of Business, Ekonominaukio 1, Otaniemi, Espoo
November 4 from 10AM to 5PM, room V2002
November 5 from 9AM to 4PM, room V2002
Registration
Registration is open until October 25, 2026.
Speaker
Professor Raghava Mukkamala, Copenhagen Business School, Denmark
Assistant Professor Sippo Rossi, Hanken School of Economics
Organizer
Professor Matti Rossi, Aalto University School of Business
Course Introduction / Description
Generative AI and large language models (LLMs) are increasingly used to support research tasks such as text coding, thematic analysis, classification, data exploration, statistical programming, interpretation, and automating multi-step analytical workflows. This course focuses on how researchers can use these technologies responsibly, critically, and reproducibly for research data analysis. The course begins with the essential fundamentals of machine learning (ML), natural language processing (NLP), and LLM concepts needed to understand what these systems can and cannot do. It then moves to responsible model use, including the distinction between proprietary hosted services and open-weight models that can be run locally or on institutional infrastructure. It pays particular attention to confidentiality, anonymization, data governance, model selection, and the practical implications of sending research data to external services.
The central part of the course is hands-on: participants use LLMs for both qualitative and quantitative data analysis. For qualitative analysis, the course covers analytical tasks such as coding, codebook development, categorization, classification, and thematic exploration. For quantitative analysis, LLMs are treated as assistants for planning analyses, generating and explaining code, interpreting outputs, and checking analytical choices; numerical results are verified through appropriate computational tools rather than accepted from the language model alone. The course concludes with advanced analytical workflows, including retrieval-augmented generation (RAG), introductory fine-tuning and domain adaptation, and AI agents or agentic workflows that can coordinate multi-step research tasks. Across all sessions, the course emphasizes reliability, hallucination, bias, transparency, reproducibility, human oversight, and ethical use.
The course is mainly designed for PhD students and early-career faculty who want to use LLMs for qualitative and quantitative data analyses. It also includes hands-on exercises on these topics using either closed-source models such as ChatGPT/Gemini/Claude or open-source models like Qwen, Llama, and so on. The PhD students are expected to have familiarity with some traditional data analysis (e.g. qualitative/quantitaive/computational).
The following is the course outline:
- The course starts with some fundamental concepts of machine learning (ML) and NLP.
- Second, it presents high-level architectures of deep learning, generative models, and LLMs, and explains why LLMs like ChatGPT have achieved so many analytical capabilities.
- Third, it presents concepts like AI Agents and Agentic workflows and how they can automate processes in organizations.
- Fourth, it will provide hands-on experience using LLMs for qualitative and quantitative data analysis, and highlight the pitfalls and limitations we need to be aware of when performing data analysis with LLMs.
Course Learning Outcomes
After completing this course, the participants should be able to:
1. Demonstrate a fundamental understanding of LLMs and how they can be used for analyzing data.
2. Design and conduct LLM-assisted qualitative data analysis workflows, including coding, categorization, classification, codebook development, and thematic exploration.
3. Use LLMs to support quantitative data analysis through computational workflows, including analysis planning, code generation, data exploration, interpretation, and result validation.
4. Describe the key challenges and opportunities, including issues related to reliability, hallucination, and ethical considerations in using Generative AI and LLMs for data analysis.
Prerequisites
Familiarity with qualitative/quantitative data analysis and experience in using LLMs with various types of prompts.
Pedagogy
Face-to-face teaching.
Program
Session Time Topic & Objective Study Material
Day-01: Wednesday, 2026-11-04
00 10:00 – 10:15 Course Introduction and Practicalities
01 10:15-12:00 Fundamentals of Machine Learning and Natural Language Processing (NLP) · Types of Machine Learning, Performance Measure · Text Classification, Topic Modeling, and others Slides, articles and other reading materials
Lunch Break
02 13:00 – 15:00 Introduction to Generative AI and Large Language Models (LLMs): transformers architecture, attention mechanism, prompt engineering Slides, articles and other reading materials
03 15:15 - 17:00 Hands-on session on Qualitative Data Analysis using LLMs
Day-02: Thursday, 2026-11-05
04 09:00 - 10:30 Fine-tuning LLMs for specific applications, hosted proprietary vs. open-weight/local models, anonymization and running LLMs locally on Confidential data Slides, articles and other reading materials
05 10:45 – 12:00 AI Agents and Agentic frameworks and Agentic workflows, building a RAG applications Slides, articles and other reading materials
Lunch Break
05 13:00 – 15:00 Hands-on session on Quantitative Data Analysis using LLMs Slides, articles and other reading materials
06 15:15 – 16:00 Wrap-up: Discussion about exam projects! Feedback and reflections on the course
Evaluation Criteria
Component Individual/Group Weightage
Final project Individual/group 100%
Credit points
Doctoral students participating in the seminar can obtain 2 credit points. This requires participating and completing the assignments.
Registration fee
This seminar is free-of-charge for Inforte.fi member organization's staff and their PhD students. For others the participation fee is 400 €. The participation fee includes access to the event and the event materials. Lunch and dinner are not included.





