Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
Training SOTA Vietnamese voice AI without selling your house for data labelling
Learn how to train Vietnamese STT and TTS models without expensive data labeling. This talk covers the data pipeline, successes, and failures with live models and real numbers.
Building state-of-the-art speech models for a low-resource language usually means an eye-watering data labelling bill. This demo walks through how we trained competitive Vietnamese STT and TTS without one - the data pipeline, what worked, and what didn’t. Live models, real audio, real numbers. The approach generalises to any under-resourced language.
Blaze provides AI voice solutions for ASEAN, including TTS, STT, and LLM-powered assistants.
- Speech-to-Text (STT)Speech-to-Text (STT) converts spoken language into written text, enabling voice-powered applications and enhanced data analysis.Speech-to-Text (STT) technology, often powered by machine learning models like those found in Amazon Comprehend Call Analytics, transforms audio input into accurate, searchable text. This capability is crucial for a wide range of applications, from transcribing customer service calls for sentiment analysis and keyword extraction to enabling voice commands in smart devices and accessibility features. For instance, businesses leverage STT to analyze vast amounts of spoken data, identifying trends, improving agent performance, and ensuring compliance. Developers integrate STT APIs (Application Programming Interfaces) into their platforms, allowing users to dictate emails, control software with voice, or generate captions for video content automatically.
- Text-to-Speech (TTS)Text-to-Speech (TTS) converts written text into natural-sounding human speech, enabling audio output for diverse applications.Text-to-Speech (TTS) technology synthesizes written language into spoken audio, mimicking human voices with impressive accuracy. Solutions like Amazon Polly and Google Cloud Text-to-Speech utilize deep learning to generate lifelike speech across multiple languages and dialects. This enables a wide range of applications: screen readers for accessibility, voice assistants (e.g., Siri, Alexa), audiobooks, and IVR systems. Developers integrate TTS via APIs, specifying parameters like voice, language, and speaking rate to customize the output. The technology continuously evolves, with advancements focusing on more natural prosody, emotional expression, and personalized voice generation.
- Vietnamese voice AIFPT.AI Voice (Text-to-Speech) delivers hyper-realistic Vietnamese AI voices with regional accents, achieving 98% human-like quality and reducing audio production costs by up to 90%.FPT.AI Voice, a leading Vietnamese voice AI technology, provides advanced Text-to-Speech (TTS) capabilities. It leverages Natural Language Processing (NLP) and deep learning to synthesize natural, context-aware speech with diverse voice options (male/female) and regional accents (Northern, Central, Southern Vietnamese). This platform boasts a 98% hyper-realistic voice quality, significantly reducing audio production costs by up to 90%. FPT.AI Voice offers seamless integration via APIs for various applications, from chatbots to customer service call centers, and includes features like the popular 'Ban Mai Voice' for engaging content creation. Other key players in Vietnamese voice AI include Zalo AI, with its Kiki voice assistant and features like dictation and voice message transcription, and VinAI, which focuses on AI research and applications.
- Data labeling pipelineAutomate and streamline the annotation of raw data for machine learning model training.A data labeling pipeline is a structured workflow for efficiently and accurately annotating large datasets. It integrates various tools and processes (e.g., human annotators, active learning, programmatic labeling) to transform raw data (images, text, audio) into labeled examples for supervised machine learning. This pipeline ensures data quality, reduces manual effort, and accelerates model development cycles, directly impacting the performance of AI applications like computer vision and natural language processing.
- Speech modelsSpeech models enhance AI's ability to accurately transcribe and understand spoken language, adapting to diverse acoustic environments and linguistic nuances.Speech models are the backbone of advanced speech-to-text (STT) and text-to-speech (TTS) systems, enabling machines to process human language with increasing accuracy. These models, often built using deep learning techniques like recurrent neural networks (RNNs) and transformers, are trained on massive datasets of audio and corresponding text. This training allows them to learn patterns in phonetics, acoustics, and language, leading to robust performance across various accents, speaking styles, and noise conditions. For example, custom speech models can be tailored for specific industries (e.g., medical, legal) to recognize specialized terminology, significantly improving transcription accuracy beyond general-purpose models. Major players like Google, Amazon (with AWS), and IBM (with Watson) continuously refine their speech models, pushing the boundaries of natural language understanding and generation.
Related talks
More from the community
VoiceReplay
Ho Chi Minh City
Discover how AI transforms text into lifelike Vietnamese audio for content creation. This talk explores advanced voice technology…
Zero shot voice cloning vs fine-tuning
San Francisco
From AI Agent Demo to Enterprise Reality Usecases
Ho Chi Minh City
Learn how GreenNode transitioned AI agent demos to enterprise reality with practical in-house use cases like Project Manager…
How I built a Multi-Agent system from scratch
Da Nang
Discover how to build a multi-agent system with context enhancement, shrinking, and unique tool result handling for effective…
qode.world: Multi-lingual AI Interviewer
Ho Chi Minh City
See how a multi-lingual AI interviewer handles orchestration, context, and adaptability in real-time, with a live trace of…
UnaMentis: A mobile, AI, voice first learning platform
Portland
Explore UnaMentis, a mobile AI voice learning platform. This talk details aggressive on-device voice model use for low-latency,…
Compose Email
Loading recent emails...