arXiv · Computation and Language · 13 Aug 2026 · paper
This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal of this work is to addr…
arXiv · Computation and Language · 13 Aug 2026 · paper
Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured…
arXiv · Computation and Language · 13 Aug 2026 · paper
We present RT-SEMamba, a fully causal speech enhancement (SE) model built upon causal time-frequency Mamba blocks. Unlike Transformer-based architectures that rely on a growing key-value ca…
arXiv · Computation and Language · 13 Aug 2026 · paper
Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important for aligning these…
arXiv · Computation and Language · 13 Aug 2026 · paper
Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt tra…
arXiv · Computation and Language · 13 Aug 2026 · paper
Audio transcription is a critical bottleneck in language documentation. While multilingual Automatic Speech Recognition (ASR) models like Whisper offer solutions, field linguists often lack…
arXiv · Computation and Language · 13 Aug 2026 · paper
Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages. However, selecting donors remains c…
arXiv · Machine Learning · 13 Aug 2026 · paper
Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing. However, most existing systems are designed for specific tas…
arXiv · Computation and Language · 13 Aug 2026 · paper
Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to…
arXiv · Artificial Intelligence · 13 Aug 2026 · paper
The Relative Transfer Matrix (ReTM), recently introduced as a generalization of the relative transfer function for multiple receivers and sources, shows promising performance when applied t…
arXiv · Machine Learning · 13 Aug 2026 · paper
Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual…
TechCrunch · 12 Aug 2026 · opinione
AI notetaking hardware has taken off over the past couple of years, with credit-card-sized devices, pendants, pins, and even transcribing earbuds all promising to capture your meetings and…
TechCrunch · 12 Aug 2026 · opinione
AI notetaking hardware has taken off over the past couple of years, with credit-card-sized devices, pendants, pins, and even transcribing earbuds all promising to capture your meetings and…
TechCrunch · 11 Aug 2026 · notizia
Google also shared numbers of how people are actually using the chatbot, with 63% of Gemini users talking directly to the assistant using the voice feature. Plus, Gemini now generates more…
Hugging Face · 10 Aug 2026
OpenAI · 03 Aug 2026 · comunicato aziendale
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.
OpenAI · 22 Jul 2026 · release software
Introducing OpenAI Presence, a proven enterprise AI agent platform that helps organizations deploy trusted voice and chat agents for customer and internal workflows.
OpenAI · 16 Jul 2026 · comunicato aziendale
Cars24 uses OpenAI-powered voice and chat agents to handle 1M+ monthly conversation minutes, recover 12% of lost leads, and bring agentic workflows to teams across the company.
Hugging Face · 15 Jul 2026 · release software
OpenAI · 10 Jul 2026 · comunicato aziendale
How Deutsche Telekom is becoming an AI-native telco with OpenAI-transforming customer service, employee workflows, network operations, and the future of voice.
OpenAI · 08 Jul 2026 · release software
A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice.
Berkeley AI Research · 01 Jul 2026 · paper
Congratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026! This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity…
Hugging Face · 01 Jul 2026
NVIDIA Developer · 09 Jun 2026
Training a speech AI model to correctly recognize or synthesize clinical terminology is surprisingly difficult. Drug names like Acetaminophen, Amlodipine,...
OpenAI · 09 Jun 2026 · comunicato aziendale
How Notion uses Codex to one-shot specs, build AI Voice Input for the web, and multiply engineering power across small teams.
OpenAI · 07 May 2026 · comunicato aziendale
Parloa leverages OpenAI models to power scalable, voice-driven AI customer service agents, enabling enterprises to design, simulate, and deploy reliable, real-time interactions.
OpenAI · 07 May 2026 · comunicato aziendale
Explore new realtime voice models in the OpenAI API that can reason, translate, and transcribe speech, enabling more natural and intelligent voice experiences.
OpenAI · 06 May 2026 · comunicato aziendale
Uber uses OpenAI to power AI assistants and voice features that help drivers earn smarter and riders book faster across a global real-time marketplace.
OpenAI · 04 May 2026 · comunicato aziendale
How OpenAI rebuilt its WebRTC stack to power real-time Voice AI with low latency, global scale, and seamless conversational turn-taking.
Hugging Face · 28 Apr 2026 · release software
Hugging Face · 24 Mar 2026
OpenAI · 20 Jan 2026 · comunicato aziendale
ServiceNow expands access to OpenAI frontier models to power AI-driven enterprise workflows, summarization, search, and voice across the ServiceNow Platform.
VentureBeat · 16 Jan 2026 · paper
Alfred Wahlforss was running out of options. His startup, Listen Labs , needed to hire over 100 engineers, but competing against Mark Zuckerberg's $100 million offers seemed impossible. So…
OpenAI · 07 Jan 2026 · tutorial
Tolan built a voice-first AI companion with GPT-5.1, combining low-latency responses, real-time context reconstruction, and memory-driven personalities for natural conversations.
Hugging Face · 28 Oct 2025
OpenAI · 30 Sep 2025 · comunicato aziendale
Sora 2 is our new state of the art video and audio generation model. Building on the foundation of Sora, this new model introduces capabilities that have been difficult for prior video mode…
OpenAI · 28 Aug 2025 · release software
We’re releasing a more advanced speech-to-speech model and new API capabilities including MCP server support, image input, and SIP phone calling support.
OpenAI · 17 Jul 2025 · comunicato aziendale
Invideo AI uses OpenAI’s GPT-4.1, gpt-image-1, and text-to-speech models to transform creative ideas into professional videos in minutes.
OpenAI · 26 Jun 2025 · comunicato aziendale
Retell AI is transforming the call center with AI voice automation powered by GPT-4o and GPT-4.1. Its no-code platform enables businesses to launch natural, real-time voice agents that cut…
Hugging Face · 09 Apr 2025
OpenAI · 20 Mar 2025 · release software
For the first time, developers can also instruct the text-to-speech model to speak in a specific way—for example, “talk like a sympathetic customer service agent”—unlocking a new level of c…
Hugging Face · 20 Dec 2024
Hugging Face · 22 Oct 2024
OpenAI · 01 Oct 2024 · release software
Developers can now build fast speech-to-speech experiences into their applications
OpenAI · 20 Jun 2024 · comunicato aziendale
Diffusion models have significantly advanced the fields of image, audio, and video generation, but they depend on an iterative sampling process that causes slow generation.
OpenAI · 07 Jun 2024 · comunicato aziendale
Exploring the technology behind our text-to-speech model.
OpenAI · 19 May 2024 · comunicato aziendale
How the voices for ChatGPT were chosen We worked with industry-leading casting and directing professionals to narrow down over 400 submissions before selecting the 5 voices.
OpenAI · 13 May 2024 · release software
We’re announcing GPT-4 Omni, our new flagship model which can reason across audio, vision, and text in real time.
OpenAI · 29 Mar 2024 · recensione
We’re sharing lessons from a small scale preview of Voice Engine, a model for creating custom voices.
Hugging Face · 27 Feb 2024
Hugging Face · 30 Aug 2023
Hugging Face · 02 Jun 2023
OpenAI · 18 May 2023 · release software
The ChatGPT app syncs your conversations, supports voice input, and brings our latest model improvements to your fingertips.
Hugging Face · 08 Feb 2023
Hugging Face · 15 Dec 2022 · tutorial
OpenAI · 21 Sep 2022 · release software
We’ve trained and are open-sourcing a neural net called Whisper that approaches human level robustness and accuracy on English speech recognition.
Hugging Face · 28 Jul 2022 · release software
OpenAI · 14 Jul 2022 · recensione
As part of our DALL·E 2 research preview, more than 3,000 artists from more than 118 countries have incorporated DALL·E into their creative workflows. The artists in our early access group…
Hugging Face · 01 Feb 2022
OpenAI · 30 Apr 2020 · release software
We’re introducing Jukebox, a neural net that generates music, including rudimentary singing, as raw audio in a variety of genres and artist styles. We’re releasing the model weights and cod…