Text Analytics

Learning essential techniques, algorithms, and models used in natural language processing. Understanding of the architectures of typical text analytics applications and of libraries for building them. Expertise in design, implementation, and evaluation of applications that exploit analysis, interpretation, and transformation of texts.

Programme Data Science and Business Informatics
Degree MSc
Institution University of Pisa
Academic year 2026/2027
Year of course 2
Period Semester 1 (15 Sep – 20 Dec 2026)
Credits 6 CFU
Language English
Scientific sector INF/01
Official page course catalogue
Teams channel join the channel

Office Hours

Book a slot here, or write me an email or Teams message.


Lectures

Date Lecture Slides Code Sources
18/09/2026 Course Introduction · Course Logistics
· Introduction to Text Analytics
· Notebook: Python Basics  
20/09/2026 Text Processing and Linguistic Analysis (1) · Text Processing and Linguistic Analysis   · J&M Ch. 2.1-2.8
· J&M Ch. 5.1
· J&M Ch. 18.1-18.3, 18.6, 18.7
· J&M Ch. 19.1, 19.2
· J&M Ch. 20.1
24/09/2026 Text Processing and Linguistic Analysis (2)   · Notebook: Regular Expressions
· Notebook: Text Processing
 
28/09/2026 Probabilistic Language Models and Collocations (1) · Probabilistic Language Models and Collocations · Notebook: Probabilistic Language Models and Collocations · J&M Ch. 3
· Stanford CS109 pg. 33-61
02/10/2026 Probabilistic Language Models and Collocations (2)     · Manning & Schütze Ch. 5
· J&M App. J
· Word association norms, mutual information, and lexicography
05/10/2026 Representing Text and Information Retrieval · Representing Text and Information Retrieval · Notebook: Representing Text · J&M Ch. 5.2-5.5
· J&M Ch. 11.1.1-11.1.4
09/10/2026 Machine Learning for Text Analytics (1) · Machine Learning Learning for Text Analytics (1) · Notebook: Machine Learning Learning for Text Analytics (1) · J&M Ch. 4.1-4.4
· J&M Ch. B.1-B.6
· J&M Ch. 23.1-23.3
· Weisberg Ch. 2
· Kumar Ch. 4

Exam

Take also a look at the Course Logistics slides for the exam modalities.

The exam can be taken in one of two ways.

  Written Exam + Oral Exam Group Project + Oral Exam
Who Anyone Attending students only
First part Written exam: 2 hours, no material allowed, typically 4 open questions Group project: a paper (max 10 pages) + code, and a presentation
Oral exam Discussion of the written exam, plus questions and exercises on course topics Discussion of the project, plus individual questions on course topics

In the written modality, the grade of the written exam is the starting grade for the oral exam.

Some examples of old written exams:

Important: from January 2027, all students (including those from previous years) must follow the current course guidelines and program.


The group project

  • Who: only students who regularly attend lessons. The project replaces the written exam.
  • Task: develop a text analytics application.
  • Topic: a challenge (SemEval, EVALITA, Kaggle), a research paper, or your own proposal.
  • Groups: 1 to 3 students.
  • Proposal deadline: the project must be agreed with the teacher by October 31st (earlier is better).
  • Final submission: 7 days before the date set for a regular exam session.
  • Exam sessions: first or second winter exam session (January to February 2027).
  • Deliverables:
    • a paper reporting on the activity (max 10 pages);
    • Python code, with comments.
  • Oral exam: presentation and discussion of the project, with individual questions on the project. This part makes up 50% of the final grade. The oral also includes questions on the course topics.

1. Timeline

  1. Register your team using the group registration form. Each team chooses a captain, who handles all communication with the teacher.
  2. Send a project proposal: the captain sends it to the teacher on behalf of the team (see Section 5 for the format).
  3. Agree on the proposal during office hours with the teacher, by October 31st.
  4. Midterm update during office hours with the teacher, to check how the project is going (November/December).
  5. Write the report with your results.
  6. Final submission of code and report, 7 days before the date of a regular exam session.
  7. Oral exam: project discussion and questions on the course topics (January to February 2027 exam sessions).

2. Choosing a topic

The topics, competitions and data sources below are only examples. You are free to choose any topic, as long as it is a text analytics task and it is agreed with the teacher.

Example topics

  • Fake review detection
  • News classification
  • Toxicity detection / hate speech detection
  • Text classification, e.g., song genre or movie genre
  • Sentiment analysis / emotion detection
  • Rating prediction
  • Detecting whether a text was written by a human or generated by AI
  • Research gaps
  • Propose your own!

Scientific NLP competitions

Where to find data

If you pick a Kaggle challenge (or any task with a ready-made benchmark), keep in mind that most of the work is already done for you: the dataset is collected and cleaned, the task is defined, and there are often thousands of public notebooks solving it. For this reason, expectations on the evaluation are higher. A project that simply reproduces a public notebook is not sufficient. You must:

  • compare your approach with other existing approaches (public notebooks, leaderboard entries, published baselines);
  • explain why your results are better or worse than theirs;
  • show what your project adds beyond what is already available.

3. Project proposal

The proposal is a short document with the following sections:

  1. Group id
  2. Motivation
  3. Goal of the project
  4. Available data
  5. Implementation and evaluation (idea / tentative plan)
  6. References
Example proposal (Group X)

Motivation

  • Online reviews play a very important role in users’ choice of whether to book, buy, rent, etc.
  • Fake reviews are commonly used to gain or damage credibility.

Project goal

  • The goal of this project is to identify patterns in the text that allow classifying it according to [a certain task] (e.g., polarity, misogyny, fake news, …).

Available data

  • Available at: Kaggle, GitHub, or another link.
  • Or: we build our dataset by merging DataSetA1 [1], containing N posts, DataSetA2 [2], containing M posts, …, and by scraping this website.
  • Context: this corpus consists of truthful and deceptive hotel reviews of 20 Chicago hotels, …
  • Content: 400 truthful positive reviews from X, 400 deceptive positive reviews from Mechanical Turk, …

Implementation

  • Text classification task (e.g., sentiment, fake vs. real, truthful vs. deceptive, …).
  • Traditional ML models, CNN, RNN, Transformer.
  • Evaluation metrics: accuracy, precision, recall, F1-score.

References

  • [1] M. Ott et al. 2011. Finding Deceptive Opinion Spam by Any Stretch of the Imagination.
  • [2] …

4. Report structure

The report is at most 10 pages, excluding references and an optional appendix. The appendix is for extra material only (additional tables, hyperparameter grids, more examples): the main text must be readable without it. The page limit is strict.

Format

We suggest writing the report in LaTeX (e.g., on Overleaf) using a standard scientific template, such as the NeurIPS template or the ACL template (the reference format for NLP conferences). Both are available on Overleaf. A template is not mandatory, but whatever you use should look like a scientific paper: title, authors with student IDs and emails, abstract, numbered sections, numbered and captioned figures and tables referenced in the text, and a bibliography. Do not shrink fonts or margins to fit the page limit.

Sections
  1. Introduction. Motivation, the task you address and why it matters, and a short summary of what you did and what you found. A reader should know your main result after reading this section.

  2. Related work / background. What others have done on the same task or data: papers, shared task results, leaderboards, public notebooks. Use it to position your work, not as a textbook: you don’t need to explain how an SVM or TF-IDF works. Cite your sources.

  3. Data understanding and preprocessing. Where the data comes from and how it was collected or labeled. Corpus statistics (number of documents, document length, vocabulary size, label distribution, class imbalance) and visualizations (e.g., length histograms, most frequent terms per class). Then the preprocessing steps, each with a short justification. If labels were not written by humans (e.g., produced by a pretrained model), say so clearly and check their quality on a sample you annotate by hand.

  4. Methods. The representations and models you use, and why you chose them for this task. Focus on your choices (features, architectures, how metadata is used, how class imbalance is handled), not on generic descriptions of the algorithms.

  5. Experiments and results.
    • Setup: train/validation/test split (or cross-validation), how hyperparameters were tuned, the evaluation metrics and why they suit the task.
    • Baselines: always include at least one simple baseline (e.g., majority class, TF-IDF + logistic regression), and, where available, results from other approaches on the same data (see the note on Kaggle-style projects).
    • Quantitative results: a comparison table of all models on the test set, plus plots where they help.
    • Qualitative analysis: look at the outputs. Show examples of correct and wrong predictions, discuss what the errors have in common (error analysis), and use visualizations such as confusion matrices, embedding projections, topics or feature importances to explain the model’s behavior.
  6. Conclusions and limitations. What you found and what it means for the task. Then be honest about the limitations: what does not work, possible biases in the data or labels, threats to the validity of the evaluation, and what you would do next.

  7. References. Every dataset, paper, model and code base you used.

Common mistakes from past years

These are the most frequent problems in reports from previous editions of the course. Avoid them.

  • Going over the page limit. Most past reports were 11 to 18 pages long. Choose what matters and move the rest to the appendix.
  • Explaining the textbook. Pages spent describing how TF-IDF, Naive Bayes or an SVM work. Assume the reader knows them; explain why you used them and how.
  • One subsection per model, twice. Methods and Results each listing SVM, Random Forest, LSTM, BERT, … with a paragraph each. Present the models together and compare them in a single results table.
  • Evaluating against a model’s own labels. Using a pretrained model to label the data and then reporting accuracy against those labels as if they were ground truth. Your models are then only imitating the labeler. If you need automatic labels, validate them by hand and discuss this as a limitation.
  • No baseline. Without a simple baseline, an F1 of 0.72 says nothing.
  • Stopping at the confusion matrix. A confusion matrix shows where the model fails, not why. Read the misclassified texts and say what goes wrong.
  • Figures without comments. Every figure and table must be referenced and discussed in the text. If you have nothing to say about it, remove it.
  • Few or no references. A report with zero to three references does not have a related work section.

5. Advice

  • Set realistic expectations (task scope, computational power, …).
  • Be methodical in defining the problem, in the use of training, validation and test data, and in the choice of evaluation measures. Working with method matters much more than achieving incredible results.
  • Be dedicated. The project requires respecting deadlines, attending tutorials, holding calls and chats with your team, and many hours of self-study.
  • Be patient and responsive. Treat your colleagues with professionalism and respect.

FAQ

Why should I do the project? Because you will learn skills that are genuinely useful.

Who can participate? Students who regularly attend lessons.

Can we do the project at any time? No. The project is developed during the course and submitted and discussed in the first or second winter exam session (January to February 2027). You must agree on a project with the teacher by October 31st.

How large can a group be? From 1 to 3 students. Register using the group registration form.

Who talks to the teacher? The team captain handles all communication with the teacher on behalf of the group.

When is the submission deadline? 7 days before the date set for the regular exam session in which you want to take the oral.

What do we submit? A paper reporting on the activity (max 10 pages) and the Python code, with comments.

How much does the project count? The presentation and discussion of the project, together with individual questions on the project, make up 50% of the final grade. The rest comes from questions on the course topics during the oral.

What happens if the project is not sufficient?

  • In the first winter exam session: you can retry in the second session, fixing or redoing the insufficient parts.
  • In the second winter exam session: you will have to take the written exam.

What happens if the oral is not sufficient?

  • In the first winter exam session: you can retry the oral in the second winter exam session, keeping the project.
  • In the second winter exam session: you will have to take the written exam.

Can I still choose the written exam? Yes. Anyone can take the written + oral exam instead.