K-INFO
HU
EN
Login

Speech Information Systems

Beszédinformációs rendszerek
A tantárgyleírás hatályossága
Hatályosság kezdete:
2026. March 21.
Hatályosság vége:
Subject name (Hungarian, English)
Beszédinformációs rendszerek
Speech Information Systems
Subject code BMEVITMAD02
Subject type
Training Level
Course types and hours (weekly/semester)
Course type lecture tutorial laboratory
hours (weekly) 2 0 2
type (linked/independent) derived course
Assessment type félévközi érdemjegy
Credits 5
Subject coordinator
DR. Németh Géza
position: egyetemi tanár
Responsible department
Távközlési és Mesterséges Intelligencia Tanszék
Faculty Villamosmérnöki és Informatikai Kar
Subject website http://smartlab.tmit.bme.hu/oktatas-beszedinformacios-rendszerek
Primary curriculum type
Direct prerequisites – Strong prerequisite none
Direct prerequisites – Weak prerequisite none
Direct prerequisites – Parallel prerequisite none
Direct prerequisites – Milestone prerequisite none
Direct prerequisites – Exclusion none

Objectives

Programme

1. Introduction
Why is speech technology important? What are the main components of a speech information system (e.g. intelligent personal assistant)?

Language, speech and text in human communication. The elements of the natural speech chain and how they work. Basic concepts of human speech production, speech perception and understanding. The main features of the acoustic structure of speech. Levels of speech, redundancy, additional information carried.

2, 3, Elementary signal processing, speech coding and compression
The role of speech coding in digital speech storage and in infocommunication network systems. Understanding the elementary signal processing steps required for the transmission, storage and analysis of speech, taking into account the specific properties of speech. Distinguish between speech/silence and other acoustic signals. Basic methods of speech coding, familiarisation with current popular formats. The impact of coding on other speech technologies. Classification of coded speech (intelligibility, naturalness).

4, Speech response systems basics
Basic concepts of machine generated speech (fixed, unlimited and mixed vocabulary). Design aspects and implementation steps for a fixed vocabulary acoustic database. Reasons for designing mixed systems and possible solutions. High fidelity prosody modification algorithms. Construction and basic classes of text-to-speech and concept-to-speech systems.

5, 6 Importance of speech and text databases, Text processing techniques
Description, design and processing methods of databases. The role of the acoustic environment. Understanding the phases of recognition creation. Databases, Vocabulary, Automatic extension of databases, adaptivity. The role of prosody. Design of multilingual systems. Developer environments and tools. Creating databases optimised for machine learning. Punctuation handling, text pre- and post-processing.


7, 8, Advanced speech response systems
Unified text representation, text analysis and transformation tasks and related databases. Design aspects and methods of construction of acoustic databases with unlimited vocabulary. Design of a speech response corpus. Database creation, modification and their algorithms. The importance and implementation of prosody (pitch, volume, rhythm variation). Multi-voice systems and automatic voice conversion. Multilingual systems. Language detection, accentuation. Unified voice annotation systems. Development environments. Algorithms for automated implementation of systems, machine learning based solutions.

9, Speech-based classification
The concept and uses of speech classification. Basic concepts of voice-based speaker identification, speaker identification and verification, UBM, likelihood-ratio framework. Voice activity detection. Introduction to the mathematical basis of the methods used. Machine learning based methods. Use of speaker identification in practice.

10, 11 Speech recognition
Basic concepts and basic architectures of speech recognition. Reference-based pattern matching. Basic equation of MAP for speech recognition. Application of Hidden Markov Models (HMM) in speech-to-text conversion. Acoustic, pronunciation and language modelling, knowledge source integration in the Weighted Finite State Transducer (WFST) framework. Hybrid HMM-neural and end-to-end neural approaches for speech recognition. CTC (Connectionist Temporal Classification) learning. Advanced neural architectures for speech recognition. Self-supervised and unsupervised speech-to-text conversion techniques. Tailoring models for application, fine-tuning. Restoring text to written form. Offline and online speech recognition, dictation process.

12, Design and implementation steps for speech information systems.
Typical application environments, dominant application systems (e.g. customer service automation, healthcare, rehabilitation). The concept of corporate acoustic image and methods to ensure a high quality company image.

13, 14 Application of speech functions in information systems
Basic concepts of speech-based dialogue systems. System driven, user driven and mixed initiative systems. DTMF and speech recognition based control in speech response systems. Uni- and multimodal systems. Modality conversion and its role in global personal communication systems.  

Lab topics:
Lab 1: Basic speech acoustics, Energy measurement, modification, spectrum generation, analysis, F0 measurement
Lab 2: Quantization, Sampling, Speech Compression, Speech Editing
Lab 3: Speech synthesis
Lab 4: Speaker identification
Lab 5: Speech recognition
Lab 6: Dialogue systems  (e.g. chatbot) 

Human information processing and communication is based on the natural speech chain (human speaker - air - listening person). Speech information systems integrate artificial information technology implementations of one or more elements of the natural speech chain (e.g. speech recognition, speech synthesis, etc.) into the processes of information collection, storage, processing and/or access. Nowadays, large-scale, increasingly integrated and automated speech information systems have appeared in many practical applications (e.g. automated speech functions in smartphones, TVs, tablets, call centres, tele-banking such as Apple Siri assistant, Google Voice Search, dictation systems, speech and text analytics, machine translation and real-time speech-to-speech interpretation). The objective of the course is to introduce the artificial implementation of the elements of the speech chain and to describe the procedures of speech-controlled and/or speech-responsive information systems that are speech-specific. The course uses practical examples to introduce the theoretical and practical knowledge required to design speech information systems, the main elements of speech technology tools for automation, their basic operating principles and specification characteristics. Students who successfully complete the course will be able to: (K1) review the basic system elements necessary for the design of speech information systems or IT systems that incorporate speech technology, (K2) develop specifications for the design of speech information systems or IT systems incorporating speech technology, (K3) design and implement test procedures for the development of speech information systems or IT systems incorporating speech technology, (K4) perform system integration tasks for the development of speech information systems or IT systems incorporating speech technology.

Learning outcomes

Ez a tantárgy a KKK rendeletben meghatározott, következő kompetenciák fejlesztését szolgálja:

Knowledge

No learning outcomes recorded.

Skills

No learning outcomes recorded.

Attitudes

No learning outcomes recorded.

Autonomy and responsibility

No learning outcomes recorded.

Oktatási módszertan

Lecture each week, 4-hour lab session every other week.

Tanulástámogató anyagok

Online források
Recommended literature:; - G. Németh, G. Olaszy: A magyar beszéd, Akadémiai Kiadó, 2010,; Available for download: http://smartlab.tmit.bme.hu/kf-letoltheto-konyvek#magyarbeszed; - Speech Recognition A Complete Guide - 2020 Edition, 5STARCooks, 2021, ISBN-13 : 978-1867335153, ISBN-13 : 978-1867335153; - D. Gardner-Bonneau: Human Factors and Voice Interactive Systems, Kluwer, 1999; - NVIDIA Nemo: https://docs.nvidia.com/deeplearning/

Recommended preliminary knowledge for completing the subject

Knowledge type competencies
(azon előzetes ismeretek összessége, amelyek megléte nem kötelező, de a tantárgy eredményes teljesítését nagyban elősegíti)
Basics of Probability Calculations
Skill type competencies
(azon előzetes képességek és készségek összessége, amelyek megléte nem kötelező, de a tantárgy eredményes teljesítését nagyban elősegíti)
nincs
Recommended (non-compulsory) preliminary competencies
(azon ajánlott (nem kötelező) előzetesen megszerzendő kompetenciák összessége, amelyek jelentősen hozzájárulnak a tantárgy eredményes teljesítéséhez)
Basics of Probability Calculations
General rules
Requirements: During term time a minor test during each lab. In case of an unsuccessful minor test, the lab cannot be completed. A report must be submitted at the end of the lab. Its acceptance is a binary decision. In the case of a failed lab, it must be repeated. The credit for the course is awarded to the student who fulfils all the following conditions. At least 5 labs are successfully completed. If the lab requirement if fulfilled, the major test result reached at least 40%. If the lab requirement if fulfilled the course grade is determined by the major test results: Additional possibilities: In exam session 1 lab can be repeated. The major test can be repeated once.
Assessment methods
In-term assessments

No detailed assessments provided.

Weight of in-term assessments

No weights provided.

Exam-period assessments

No detailed assessments provided.

Weight of exam elements

No weights provided.

Grade calculation

No grade thresholds provided.

Attendance requirements

No attendance requirements provided.

Rules for retake and resubmission

Not provided.

Short description

Not provided.

Detailed description
IMSc program: We want to reward outstanding solutions f the small tests and large test with IMSc points. In addition, for interested students, an individual assignment will be issued, which can also be used to obtain IMSc points. IMSc points: If the overall average of the top five small tests reaches 85%, the student will receive 10 IMSc points. The number of IMSc points awarded in the major test is equal to the number of IMSc points the student scores above the 85% threshold. A maximum of 10 IMSc points will be awarded for solving an individual assignment in the subject. The total IMSc points for a student may not exceed 25.
Recommended courses

Not provided.

Workload to complete the subject

No workload breakdown provided.

Validity of subject requirements
Requirements valid from:
Requirements valid until:
Curriculum placement

No curriculum placements recorded for this subject version.