Vocal Isolation
Extract clean vocal tracks from any music. Separate Vocals mode splits audio into isolated vocals and instrumentals with high-quality output, perfect for karaoke, remixing, or vocal analysis.

Transparent pricing with no hidden fees. Pay as you go.
separate-vocals
stem-split
upload-and-separate-vocals| Model | Spec | PoYo price | Official price | You save |
|---|---|---|---|---|
separate-vocals | Standard generation | $0.075/gen 15 credits /gen | - | - |
stem-split | Standard generation | $0.350/gen 70 credits /gen | - | - |
upload-and-separate-vocals | Standard generation | $0.050/gen 10 credits /gen | - | - |
* Actual fees are based on the final output.
Complete guide to using Vocal Remover API - AI-Powered Audio Separation on PoYo
Extract clean vocal tracks from any music. Separate Vocals mode splits audio into isolated vocals and instrumentals with high-quality output, perfect for karaoke, remixing, or vocal analysis.
Decompose audio into individual instrument stems. Stem Split mode separates drums, bass, vocals, and other instruments into isolated tracks for professional music production and remixing.
Choose Separate Vocals (15 credits) for simple vocal/instrumental separation, or Stem Split (70 credits) for advanced multi-track decomposition. Select the right mode for your needs and budget.
Professional-grade audio separation with minimal artifacts. Advanced AI models ensure clean separation while preserving the original audio quality and fidelity.
Support for common audio formats including MP3, WAV, FLAC, and more. Process any audio file and receive separated tracks in high-quality output format.
Quick audio separation with optimized AI processing. Webhook callbacks notify you instantly when your separated tracks are ready for download.
Extract stems for remixing, mashups, and music production. Isolate vocals, drums, bass, and instruments to create new arrangements and creative compositions.
Generate karaoke tracks by removing vocals from songs. Create instrumental backing tracks for singing practice, karaoke apps, and vocal training platforms.
Isolate voice tracks from background music for cleaner audio editing. Separate audio elements for more precise control in podcast and video production.
Create learning materials by isolating individual instruments. Help students study specific parts, practice along with isolated tracks, and understand musical arrangements.
Separate audio tracks for more flexible sound design and mixing. Extract specific audio elements for re-editing, dubbing, and creative sound design in post-production.
Isolate audio sources for detailed analysis and study. Perfect for music information retrieval, audio forensics, and academic research on musical composition.
1. Get Your API Key: Create a free PoYo account and generate your API key from the dashboard. Get API Key →
2. Choose Your Model: Select from three models based on your needs - Separate Vocals for generated music, Upload and Separate Vocals for your own audio files, or Stem Split for multi-track decomposition.
3. Make API Request: Send a POST request to the appropriate endpoint with required parameters. View documentation for each model:
4. Get Results: Receive a task_id and poll for results, or set up a webhook callback URL to be notified when separation completes with download URLs for each isolated track.
Separate Vocals: $0.075 per separation (15 credits) - Split audio into vocals and instrumentals
Stem Split: $0.35 per separation (70 credits) - Decompose into multiple stems (drums, bass, vocals, etc.)
Features Included: High-quality audio separation, multiple format support, webhook callbacks, professional-grade output
Enterprise: Volume discounts available for high-volume usage. Priority processing and dedicated support. Contact sales →
Choose between two modes based on your needs: simple vocal separation at $0.075 or comprehensive stem splitting at $0.35. Pay only for what you use with transparent credit-based pricing.
Access Suno's state-of-the-art audio separation technology through PoYo's unified API. Professional-quality output with industry-leading AI models.
Quick processing times with 99.9% uptime SLA. Webhook notifications ensure you're instantly alerted when your separated tracks are ready.
Studio-grade audio separation with minimal artifacts and bleed. Suitable for professional music production, broadcasting, and commercial applications.
One API key connects you to Vocal Remover and 15+ other music APIs plus 500+ AI models. Simplify billing and reduce integration complexity.
Clean RESTful API with comprehensive documentation. Simple integration with any programming language. SDKs and code examples available.