深度学习系列56：使用whisper进行语音转文字

1. openai-whisper

这应该是最快的使用方式了。安装

pip install -U openai-whisper

，接着安装ffmpeg，随后就可以使用了。模型清单如下：
在这里插入图片描述

第一种方式，使用命令行：

whisper japanese.wav --language Japanese  --model medium

另一种方式，使用python调用：

import whisper
model = whisper.load_model("base")
result = model.transcribe("audio.mp3",initial_prompt='以下是普通话的句子。')
print(result["text"])

2. faster-whisper

安装也一样：

pip install -U faster-whisper

，速度对比：
在这里插入图片描述

3. whisper-jax

在GPU上的加速版本
首先安装库：
pip install jax jaxlib git+https://github.com/sanchit-gandhi/whisper-jax.git datasets soundfile librosa

调用代码为：

from whisper_jax import FlaxWhisperPipline
import jax.numpy as jnp
pipeline = FlaxWhisperPipline("openai/whisper-tiny", dtype=jnp.bfloat16, batch_size=16)
%time text = pipeline('test.mp3')

4. whisper-openvino

在intel系列的cpu上加速的版本：
安装库：pip install git+https://github.com/zhuzilin/whisper-openvino.git
调用方法：whisper carmack.mp3 --model tiny.en --beam_size 3

标签： whisper

本文转载自: https://blog.csdn.net/kittyzc/article/details/135916306
版权归原作者 IE06 所有，如有侵权，请联系我们删除。

深度学习系列56：使用whisper进行语音转文字

1. openai-whisper

2. faster-whisper

3. whisper-jax

4. whisper-openvino

发表评论

“深度学习系列56：使用whisper进行语音转文字”的评论:

关于作者

overfit同步小助手

相关阅读

文章导航