Subject » BMEVITMMSMA002-00
Data Science on Structured and Textual Data
Adattudomány alkalmazása strukturált és szöveges adatokon
A tantárgyleírás hatályossága
Hatályosság kezdete:
—
Hatályosság vége:
—
| Subject name (Hungarian, English) |
Adattudomány alkalmazása strukturált és szöveges adatokon
Data Science on Structured and Textual Data
|
||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Subject code | BMEVITMMSMA002-00 | ||||||||||||
| Subject type | — | ||||||||||||
| Training Level | — | ||||||||||||
| Course types and hours (weekly/semester) |
|
||||||||||||
| Assessment type | vizsga | ||||||||||||
| Credits | 5 | ||||||||||||
| Subject coordinator |
Szűcs Gábor, PhD
position: adjunktus
contact:
szucs.gabor@vik.bme.hu
|
||||||||||||
| Responsible department |
—
|
||||||||||||
| Faculty | |||||||||||||
| Subject website | — | ||||||||||||
| Primary curriculum type | — | ||||||||||||
| Direct prerequisites – Strong prerequisite | none | ||||||||||||
| Direct prerequisites – Weak prerequisite | none | ||||||||||||
| Direct prerequisites – Parallel prerequisite | none | ||||||||||||
| Direct prerequisites – Milestone prerequisite | none | ||||||||||||
| Direct prerequisites – Exclusion | none |
Objectives
Programme
- Time series analysis: forecasting and classification (with deep learning methods). Hierarchical time series forecasting. Independent forecasting models. Making predictions coherent. Reconciliation methods.
- Scalable time series analysis. Approximate reconciliation methods for large data sets.
- Hierarchical classification. Multiclass and multilabel classification.
- Multiclass problems with binary classifiers: Error-Correcting Output Codes. Ensemble methods.
- Recommender systems. Content-based filtering, Collaborative filtering methods. Graph-based models for recommender systems.
- Specifics of testing recommender systems. Evaluation techniques.
- What geometry tells us about language. Practical tour of word embeddings (skip-gram, CBOW, GloVe). Cosine similarity, linear analogy offsets, and bias subspaces. Cultural bias with WEAT (Word Embedding Association Test). Domain adaptation via embedding fine-tuning. Vector drift. Analogy-based retrieval engine.
Train skip-gram on a Wikipedia subset. Run WEAT to detect and debias gender stereotypes Fine-tune on a biomedical corpus (e.g. PubMed) to measure semantic shift for ambiguous terms.
- When LLMs disagree — using ensemble uncertainty for trustworthy inference.
Majority-vote ensemble for high-stakes classification. Condorcet jury theorem.
Self-consistency sampling for calibration. Ground-truth free reliability signal
Human-in-the-loop triage via uncertainty flags.
- Train three local open-weight models (Llama, Mistral, Qwen) via Ollama. Implement plurality voting, self-consistency calibration on datasets like MedQA and GSM8K. Build a selective-prediction pipeline that trades coverage for accuracy with an explicit cost model.
- Adversarial manipulation of RAG (Retrieval-Augmented Generation) and LLM pipelines
Prompt injection. PoisonedRAG. Indirect prompt injection via external content
Defence: Detection and mitigation. Perplexity-based and query-rewriting defences.
Input classifiers, spotlighting, LLM-as-critic, instruction hierarchy. Output flip.
(Web-RAG: from raw text data to structured data)
- Build a minimal RAG system with LangChain and Chroma. Inject poisoned passages and measure attack success rate. Implement perplexity-based and query-rewriting defences and compare direct versus indirect injection against a naive and a hardened pipeline.
- Change detection in classification tasks. Continuous machine learning, Class-Incremental-Learning
Exercises:
1. Using time series reconciliation methods on real datasets
2. Classification tasks on hierarchically labeled datasets
3. Collaborative filtering in recommender systems
4. Practice on different word embeddings
5. Implement plurality voting, self-consistency calibration on datasets
6. Build a minimal RAG system with LangChain and Chroma
7. Change detection in machine learning tasks
The goal is to comprehensively understand and apply data science techniques across diverse data types, including structured, unstructured, textual, time series, and multimodal data, by data preprocessing, transformation, and modeling methods. The course enables learners to effectively analyze complex datasets and develop intelligent systems across various domains. The obtained knowledge encompasses a broad and interdisciplinary understanding of data science, including large language models.
Learning outcomes
Ez a tantárgy a KKK rendeletben meghatározott, következő kompetenciák fejlesztését szolgálja:
Knowledge
No learning outcomes recorded.
Skills
No learning outcomes recorded.
Attitudes
No learning outcomes recorded.
Autonomy and responsibility
No learning outcomes recorded.
Oktatási módszertan
lecture and practice
Tanulástámogató anyagok
Not provided.
Recommended preliminary knowledge for completing the subject
Knowledge type competencies
(azon előzetes ismeretek összessége, amelyek megléte nem kötelező, de a tantárgy eredményes teljesítését nagyban elősegíti)
statistics, data analysis, machine learning
Skill type competencies
(azon előzetes képességek és készségek összessége, amelyek megléte nem kötelező, de a tantárgy eredményes teljesítését nagyban elősegíti)
nincs
Recommended (non-compulsory) preliminary competencies
(azon ajánlott (nem kötelező) előzetesen megszerzendő kompetenciák összessége, amelyek jelentősen hozzájárulnak a tantárgy eredményes teljesítéséhez)
statistics, data analysis, machine learning
General rules
Requirements:
Six small homework assignments given every two weeks, of which at least 4 must be completed at an appropriate level to be signed. Small homework assignments cannot be made up.
Written exam covering the theoretical and practical material covered, linked to homework assignments. The passing grade for the exam is 50%.
Additional possibilities:
A maximum of two out of six small homework assignments may be skipped (or solving them at an inappropriate level is equivalent to skipping them).
Assessment methods
In-term assessments
No detailed assessments provided.
Weight of in-term assessments
No weights provided.
Exam-period assessments
No detailed assessments provided.
Weight of exam elements
No weights provided.
Grade calculation
No grade thresholds provided.
Attendance requirements
No attendance requirements provided.
Rules for retake and resubmission
Not provided.
Short description
Not provided.
Detailed description
Not provided.
Recommended courses
Not provided.
Workload to complete the subject
No workload breakdown provided.
Validity of subject requirements
Requirements valid from:
—
Requirements valid until:
—
Curriculum placement
No curriculum placements recorded for this subject version.