-
SoundCTM: Uniting Score-based and Consistency Models for Text-to-Sound Generation
Paper • 2405.18503 • Published • 9 -
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
Paper • 2405.20289 • Published • 11 -
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
Paper • 2406.02897 • Published • 15 -
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
Paper • 2406.03344 • Published • 20
Collections
Discover the best community collections!
Collections including paper arxiv:2601.15621
-
Qwen3-TTS Technical Report
Paper • 2601.15621 • Published • 80 -
PaperBanana: Automating Academic Illustration for AI Scientists
Paper • 2601.23265 • Published • 230 -
Moonshine: Speech Recognition for Live Transcription and Voice Commands
Paper • 2410.15608 • Published • 14 -
PersonaLive! Expressive Portrait Image Animation for Live Streaming
Paper • 2512.11253 • Published • 41
-
SoundCTM: Uniting Score-based and Consistency Models for Text-to-Sound Generation
Paper • 2405.18503 • Published • 9 -
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
Paper • 2405.20289 • Published • 11 -
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
Paper • 2406.02897 • Published • 15 -
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
Paper • 2406.03344 • Published • 20
-
Qwen3-TTS Technical Report
Paper • 2601.15621 • Published • 80 -
MemOS: A Memory OS for AI System
Paper • 2507.03724 • Published • 170 -
Self-Supervised Prompt Optimization
Paper • 2502.06855 • Published • 18 -
A decoder-only foundation model for time-series forecasting
Paper • 2310.10688 • Published • 46
-
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
Paper • 2503.04721 • Published • 4 -
Qwen3-TTS Technical Report
Paper • 2601.15621 • Published • 80 -
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Text-to-Speech • 2B • Updated • 2.43M • 2.02k -
openbmb/AgentCPM-Report
Text Generation • 8B • Updated • 266 • 305
-
SoundCTM: Uniting Score-based and Consistency Models for Text-to-Sound Generation
Paper • 2405.18503 • Published • 9 -
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
Paper • 2405.20289 • Published • 11 -
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
Paper • 2406.02897 • Published • 15 -
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
Paper • 2406.03344 • Published • 20
-
SoundCTM: Uniting Score-based and Consistency Models for Text-to-Sound Generation
Paper • 2405.18503 • Published • 9 -
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
Paper • 2405.20289 • Published • 11 -
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
Paper • 2406.02897 • Published • 15 -
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
Paper • 2406.03344 • Published • 20
-
Qwen3-TTS Technical Report
Paper • 2601.15621 • Published • 80 -
MemOS: A Memory OS for AI System
Paper • 2507.03724 • Published • 170 -
Self-Supervised Prompt Optimization
Paper • 2502.06855 • Published • 18 -
A decoder-only foundation model for time-series forecasting
Paper • 2310.10688 • Published • 46
-
Qwen3-TTS Technical Report
Paper • 2601.15621 • Published • 80 -
PaperBanana: Automating Academic Illustration for AI Scientists
Paper • 2601.23265 • Published • 230 -
Moonshine: Speech Recognition for Live Transcription and Voice Commands
Paper • 2410.15608 • Published • 14 -
PersonaLive! Expressive Portrait Image Animation for Live Streaming
Paper • 2512.11253 • Published • 41
-
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
Paper • 2503.04721 • Published • 4 -
Qwen3-TTS Technical Report
Paper • 2601.15621 • Published • 80 -
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Text-to-Speech • 2B • Updated • 2.43M • 2.02k -
openbmb/AgentCPM-Report
Text Generation • 8B • Updated • 266 • 305