K-INFO
HU
EN
Login
Subject » BMEVITMMSMA002-00

Data Science on Structured and Textual Data

Adattudomány alkalmazása strukturált és szöveges adatokon
A tantárgyleírás hatályossága
Hatályosság kezdete:
Hatályosság vége:
Subject name (Hungarian, English)
Adattudomány alkalmazása strukturált és szöveges adatokon
Data Science on Structured and Textual Data
Subject code BMEVITMMSMA002-00
Subject type
Training Level
Course types and hours (weekly/semester)
Course type lecture tutorial laboratory
hours (weekly) 2 1 0
type (linked/independent) derived course
Assessment type vizsga
Credits 5
Subject coordinator
Szűcs Gábor, PhD
position: adjunktus
Responsible department
Faculty
Subject website
Primary curriculum type
Direct prerequisites – Strong prerequisite none
Direct prerequisites – Weak prerequisite none
Direct prerequisites – Parallel prerequisite none
Direct prerequisites – Milestone prerequisite none
Direct prerequisites – Exclusion none

Objectives

Programme
- Time series analysis: forecasting and classification (with deep learning methods). Hierarchical time series forecasting. Independent forecasting models. Making predictions coherent. Reconciliation methods.
- Scalable time series analysis. Approximate reconciliation methods for large data sets.
- Hierarchical classification. Multiclass and multilabel classification.
- Multiclass problems with binary classifiers: Error-Correcting Output Codes. Ensemble methods.
- Recommender systems. Content-based filtering, Collaborative filtering methods. Graph-based models for recommender systems.
- Specifics of testing recommender systems. Evaluation techniques.
- What geometry tells us about language. Practical tour of word embeddings (skip-gram, CBOW, GloVe). Cosine similarity, linear analogy offsets, and bias subspaces. Cultural bias with WEAT (Word Embedding Association Test). Domain adaptation via embedding fine-tuning. Vector drift. Analogy-based retrieval engine.
Train skip-gram on a Wikipedia subset. Run WEAT to detect and debias gender stereotypes Fine-tune on a biomedical corpus (e.g. PubMed) to measure semantic shift for ambiguous terms.
- When LLMs disagree — using ensemble uncertainty for trustworthy inference.
Majority-vote ensemble for high-stakes classification. Condorcet jury theorem.
Self-consistency sampling for calibration. Ground-truth free reliability signal
Human-in-the-loop triage via uncertainty flags.
- Train three local open-weight models (Llama, Mistral, Qwen) via Ollama. Implement plurality voting, self-consistency calibration on datasets like MedQA and GSM8K. Build a selective-prediction pipeline that trades coverage for accuracy with an explicit cost model.
- Adversarial manipulation of RAG (Retrieval-Augmented Generation) and LLM pipelines
Prompt injection. PoisonedRAG. Indirect prompt injection via external content
Defence: Detection and mitigation. Perplexity-based and query-rewriting defences.
Input classifiers, spotlighting, LLM-as-critic, instruction hierarchy. Output flip.
(Web-RAG: from raw text data to structured data)
- Build a minimal RAG system with LangChain and Chroma. Inject poisoned passages and measure attack success rate. Implement perplexity-based and query-rewriting defences and compare direct versus indirect injection against a naive and a hardened pipeline.
- Change detection in classification tasks. Continuous machine learning, Class-Incremental-Learning
 
Exercises:
1. Using time series reconciliation methods on real datasets
2. Classification tasks on hierarchically labeled datasets
3. Collaborative filtering in recommender systems
4. Practice on different word embeddings
5. Implement plurality voting, self-consistency calibration on datasets
6. Build a minimal RAG system with LangChain and Chroma
7. Change detection in machine learning tasks 
The goal is to comprehensively understand and apply data science techniques across diverse data types, including structured, unstructured, textual, time series, and multimodal data, by data preprocessing, transformation, and modeling methods. The course enables learners to effectively analyze complex datasets and develop intelligent systems across various domains. The obtained knowledge encompasses a broad and interdisciplinary understanding of data science, including large language models.

Learning outcomes

Ez a tantárgy a KKK rendeletben meghatározott, következő kompetenciák fejlesztését szolgálja:

Knowledge

No learning outcomes recorded.

Skills

No learning outcomes recorded.

Attitudes

No learning outcomes recorded.

Autonomy and responsibility

No learning outcomes recorded.

Oktatási módszertan

lecture and practice

Tanulástámogató anyagok

Not provided.

Recommended preliminary knowledge for completing the subject

Knowledge type competencies
(azon előzetes ismeretek összessége, amelyek megléte nem kötelező, de a tantárgy eredményes teljesítését nagyban elősegíti)
 statistics, data analysis, machine learning
Skill type competencies
(azon előzetes képességek és készségek összessége, amelyek megléte nem kötelező, de a tantárgy eredményes teljesítését nagyban elősegíti)
nincs
Recommended (non-compulsory) preliminary competencies
(azon ajánlott (nem kötelező) előzetesen megszerzendő kompetenciák összessége, amelyek jelentősen hozzájárulnak a tantárgy eredményes teljesítéséhez)
 statistics, data analysis, machine learning
General rules
Requirements: Six small homework assignments given every two weeks, of which at least 4 must be completed at an appropriate level to be signed. Small homework assignments cannot be made up. Written exam covering the theoretical and practical material covered, linked to homework assignments. The passing grade for the exam is 50%. Additional possibilities: A maximum of two out of six small homework assignments may be skipped (or solving them at an inappropriate level is equivalent to skipping them). 
Assessment methods
In-term assessments

No detailed assessments provided.

Weight of in-term assessments

No weights provided.

Exam-period assessments

No detailed assessments provided.

Weight of exam elements

No weights provided.

Grade calculation

No grade thresholds provided.

Attendance requirements

No attendance requirements provided.

Rules for retake and resubmission

Not provided.

Short description

Not provided.

Detailed description

Not provided.

Recommended courses

Not provided.

Workload to complete the subject

No workload breakdown provided.

Validity of subject requirements
Requirements valid from:
Requirements valid until:
Curriculum placement

No curriculum placements recorded for this subject version.