Found 47 results for “Speech”
OpenGPT is a tool platform for building ChatGPT applications based on APIs, supporting capabilities such as multilingual support, instant messaging, speech recognition, and natural language processing, while also providing reference application examples and open-source code.
Lovo is an AI voice generation and text-to-speech tool that supports converting text into natural speech, suitable for audio content production, voiceover, and various creative scenarios, helping reduce manual recording costs and time investment.
YouWhisper is a community Hugging Face Space for audio and video transcription, using Whisper-style speech-to-text workflows to create transcripts.
AI Voice Detector is an audio authenticity detection tool used to identify whether speech is generated by AI. Users can upload audio files for verification, making it suitable for scenarios involving evidence review, media judgment, and authenticity analysis in customer communications.
NetEase Jianwai Workbench is an AI tool for office and collaboration scenarios, providing video and livestream transcription, speech-to-text, document translation, and other functions, suitable for teams handling multimedia and text content.
Article.Audio is an online service that converts article content into spoken audio, supporting the transformation of text articles into listenable audio for convenient access to information when reading is inconvenient.
AI Majic is an AI writing tool for content creation that can generate video descriptions, tags, social media copy, article summaries, speech points, and more, helping users complete various text content more quickly.
Transkribieren is an audio transcription tool that supports uploading multiple audio formats, provides a relatively convenient speech-to-text service, and extends to use on mobile, in the browser, and in meeting scenarios.
Adobe Speech Enhancer is an AI audio enhancement tool for improving the quality of voice recordings. It can reduce background noise and highlight voices, making ordinary spoken recordings sound clearer and closer to a studio effect.
Murf AI is an AI voice generation tool that converts text into natural, lifelike human speech, suitable for creating podcasts, video voiceovers, presentation narration, and other audio content.
Sumly.AI is an AI summary tool for podcast content that uses speech-to-text technology to distill key points from episodes and help users quickly understand podcast content through short summaries.
AI Depot is a platform that aggregates various types of artificial intelligence tools, covering areas such as text analysis, speech recognition, image recognition, and predictive analytics, helping users find suitable machine learning capabilities for different types of applications.
NeuroSpell is a deep learning-based spelling and grammar auto-correction tool that supports more than 30 languages and provides capabilities such as speech-to-text, OCR error correction, and customizable terminology training.
Otter AI is a meeting recording and note-taking tool that supports real-time speech transcription, audio recording, slide capture, and automatic meeting summary generation for easier organization and review.
GistReader is an AI-powered RSS reader that offers automatic article summaries and text-to-speech features, helping users obtain information more efficiently and save reading time.
Voiceful provides game character voice generation and speech synthesis demos, and supports integration into Unity via SDK, making it suitable for development and testing scenarios that require character voice capabilities.
Wisecut is an online AI video editing tool that uses speech recognition to automatically process video content, remove pauses, generate subtitles, and add background music. It is suitable for quickly organizing talking-head, interview, and podcast videos.
Good Tape is an automatic speech-to-text tool that can quickly convert audio recordings into text, supports more than 90 languages, and is suitable for organizing interviews, meetings, and dictated content.
Free Text to Speech Generator is an online TTS tool that supports multiple languages, multiple dialects, and mixed Chinese-English reading, and can convert text into speech and export MP3 files.
Salient is an AI tool for sales teams that can be used for personalized outbound emails, automatic customer replies, and reactivating leads, while also providing employee analytics to help businesses observe workforce trends.
NaturalReader is a tool that provides AI text-to-speech services, supporting online use, mobile apps, commercial licensing, and educational scenarios, suitable for converting text content into audio for listening.
SpeechEasy provides high-quality text-to-speech services
Thekeys is an AI writing assistant tool that helps users optimize their wording without changing the original meaning. It focuses on making text more vivid, concise, and persuasive, making it suitable for writing, speeches, and everyday communication scenarios.
Voicera is an article-to-speech tool that can automatically detect content and generate a playable audio version. It supports multiple languages and voice options, making it convenient for users to access information by listening.
Krater.AI is an AI-driven content creation tool that provides a suite of tools for marketers and content creators. It offers solutions such as ad copy generation, creating stunning images, and converting audio content into written content or realistic voiceovers. The website claims to have advanced technology comparable to Jasper, Midjourney, and Writer.com. Krater.AI is designed to be user-friendly and intuitive, offering features such as image generation, copywriting, chat, speech-to-text, code, and more. The website also has a Twitter account where they share updates about their product.
AiPy is a free and open-source AI agent factory, a local version of Manus, built on large language models (LLMs) and Python capabilities. It supports local deployment to ensure data privacy and security. Through the "Python-Use" paradigm, AiPy gives AI "hands," enabling it to analyze local data, operate local applications, and execute complex tasks such as controlling phones, generating multi-voice speech, analyzing medical test reports, extracting speech from videos, and sending scheduled emails.
AnyGen is an AI office agent launched by ByteDance that improves office efficiency through voice input and AI technology. Users can press and hold the recording button to quickly convert speech into text, with support for adding photos, screenshots, and links, avoiding the tedious organization required after traditional note-taking.
AI models for transcribing and understanding speech
Deepgram is a platform that provides advanced AI speech recognition and natural language processing technology. Its core products are powerful Speech-to-Text (STT) and Text-to-Speech (TTS) APIs, enabling developers to quickly integrate voice transcription and understanding capabilities into their own applications and services.
ElevenLabs is an AI text-to-speech platform that provides realistic voice synthesis solutions for developers, creators, and enterprises. Its core products include text-to-speech (supporting 29+ languages including Chinese and 10,000+ voices), AI dubbing, voice cloning, music generation, and more.
Feishu Minutes offers intelligent meeting notes and fast AI speech-to-text transcription.
Huawen Bigan is a professional AI official document writing tool designed for government and enterprise writers, providing full-process support from drafting to finalization. It offers functions such as writing based on existing drafts, meeting minutes, and AI official document writing, and can generate notices, reports, speeches, and other official documents with one click. Huawen Bigan supports formatting and article polishing, covering a wide range of writing scenarios.
IBM Watson Text to Speech
JoyPix is an AI creation tool focused on digital humans and speech synthesis. Users can create personalized virtual avatars by uploading photos, with support for voice conversations with virtual avatars.
LangLang Voiceover is an intelligent text-to-speech tool that provides voice synthesis services. It supports more than 30 languages, including Chinese, English, German, and French, as well as more than 10 emotional styles such as happy, sad, and excited. The platform is feature-rich and easy to use, supporting SSML tags to enable advanced functions such as polyphonic character handling and multi-speaker dubbing.
Maier Meeting Notes is an application software under AISpeech that integrates real-time speech transcription, real-time translation, AI summary analysis, and other functions, mainly used in scenarios such as office meetings, students' online classes, and customer interview recordings. The software supports recording and transcription at the same time, and after recording ends, the audio and text are synchronized in real time to the PC and mobile sides.
Miaochuang (formerly "Yizhen Miaochuang") is an intelligent AI content generation platform based on the Miaochuang AIGC engine, providing creators and organizations with AI generation services including text continuation, text-to-speech, text-to-image, and image-and-text-to-video. Yizhen Miaochuang intelligently analyzes copy, assets, AI voice, subtitles, and more to quickly produce finished videos, enabling zero-threshold video creation.
Moyin Workshop is a professional AI voiceover tool with more than 800 voices and over 1,000 styles, meeting a wide range of needs from video dubbing to audiobooks. Moyin Workshop offers rich features, including speech rate adjustment, polyphonic character selection, and pause control, ensuring realistic and natural text-to-speech results. Users can easily download lossless audio files and enjoy a convenient voiceover experience.
Quickie is an AI productivity tool that provides text-to-speech, content summarization, text expansion, and other functions, and is used as a browser extension to help users quickly process information during daily browsing and writing.
SoundView is an AI video localization tool that supports video dubbing and video translation. SoundView integrates multilingual translation, speech synthesis, speech recognition, and large-model technology to simplify and accelerate the creation of product marketing videos. SoundView supports dubbing and subtitle editing in 100 languages, increasing video production efficiency by 10 times and reducing video translation costs by 90%.
Tingnao AI is an AI-powered intelligent voice assistant focused on speech-to-text and real-time recording summaries, offering audio/video transcription, real-time recording-to-text, AI summaries, chapter overview, and other features. Users can freely drag text to view audio/video progress and enjoy a convenient intelligent recording experience.
Uberduck is an open-source community for AI voice generation and synthesis. The platform offers more than 5,000 voices to help users create AI dubbing and speech, and you can even use your own custom voice clone for synthesis.
AI text-to-speech generation tool
Wacuowang is an AI content review and proofreading platform that detects content and automatically corrects errors with one click, supporting the review of text, images, audio, video, and other content formats. Wacuowang supports rapid identification of typos, punctuation errors, speech rate issues, sensitive words, and politically related information, providing highlighted prompts and correction suggestions to ensure content accuracy and rigor.
iFLYTEK Simultaneous Interpretation is a professional AI simultaneous interpretation product launched by iFLYTEK. Based on its world-leading intelligent speech and language technologies, it provides integrated simultaneous interpretation services for multiple scenarios and languages, including real-time transcription and translation, simultaneous interpretation, live subtitle display, and meeting record sharing.
iFlytek Zhiwen is an intelligent document AI assistant launched by iFLYTEK based on the Spark large model, designed to improve the creation and presentation efficiency of Word and PPT. The tool supports features such as intelligent rehearsal and AI Presenter to help users optimize the entire process from content creation to speech delivery.
Xunfei Zhizuo is a one-stop AIGC content creation platform launched by iFLYTEK, providing services such as text-to-speech and virtual digital human video production based on artificial intelligence technology. Users can easily achieve rapid generation of audio and video content and create high-quality media works without professional skills.
