Azure Speech Provider
기준일: 2026-07-26
난이도: 중급
공식 기준: Azure Speech
개요
이 페이지는 OpenClaw Azure Speech provider의 인증, 모델 ref, 온보딩 플래그, 설정 키를 공식 문서 기준으로 정리합니다.
공식 요약: Azure AI Speech text-to-speech for OpenClaw replies
model ref는 보통 provider/model 형식입니다. API key·endpoint·model id는 아래 원문 값을 그대로 쓰세요. 임의로 키나 엔드포인트를 만들지 마세요.
빠른 참조
| Detail | Value |
|---|---|
| Provider ID | azure-speech (alias: azure) |
| Website | Azure AI Speech |
| Docs | Speech REST text-to-speech |
| Auth | AZURE_SPEECH_KEY plus AZURE_SPEECH_REGION |
| Default voice | en-US-JennyNeural |
| Default file output | audio-24khz-48kbitrate-mono-mp3 |
| Default voice-note file | ogg-24khz-16bit-mono-opus |
공식 문서 기반 상세
아래는 공식 providers/azure-speech 문서를 Mintlify 컴포넌트만 정리하고 공통 제목을 한국어로 맞춘 내용입니다. 설정 키, env, model ref, CLI 플래그는 원문 그대로 유지합니다.
Azure Speech is a bundled Azure AI Speech text-to-speech provider. OpenClaw
calls the Azure Speech REST API directly with SSML, synthesizing MP3 for
standard replies, native Ogg/Opus for voice notes, and 8 kHz mulaw for
telephony channels such as Voice Call. The request sends the provider-owned
output format through the X-Microsoft-OutputFormat header.
| Detail | Value |
|---|---|
| Provider ID | azure-speech (alias: azure) |
| Website | Azure AI Speech |
| Docs | Speech REST text-to-speech |
| Auth | AZURE_SPEECH_KEY plus AZURE_SPEECH_REGION |
| Default voice | en-US-JennyNeural |
| Default file output | audio-24khz-48kbitrate-mono-mp3 |
| Default voice-note file | ogg-24khz-16bit-mono-opus |
시작하기
Create an Azure Speech resource
In the Azure portal, create a Speech resource. Copy **KEY 1** from
Resource Management > Keys and Endpoint, and copy the resource location
such as `eastus`.
```
AZURE_SPEECH_KEY=<speech-resource-key>
AZURE_SPEECH_REGION=eastus
```
Select Azure Speech in tts
```json5
{
tts: {
auto: "always",
provider: "azure-speech",
providers: {
"azure-speech": {
voice: "en-US-JennyNeural",
lang: "en-US",
},
},
},
}
```
Send a message
Send a reply through any connected channel. OpenClaw synthesizes the audio
with Azure Speech and delivers MP3 for standard audio, or Ogg/Opus when
the channel expects a voice note.
Configuration options
All options live under tts.providers["azure-speech"].
| Option | Description |
|---|---|
apiKey |
Azure Speech resource key. Falls back to AZURE_SPEECH_KEY, AZURE_SPEECH_API_KEY, or SPEECH_KEY. |
region |
Azure Speech resource region. Falls back to AZURE_SPEECH_REGION or SPEECH_REGION. |
endpoint |
Optional Azure Speech endpoint override. Falls back to trusted AZURE_SPEECH_ENDPOINT. |
baseUrl |
Optional Azure Speech base URL override. |
voice |
Azure voice ShortName (default en-US-JennyNeural). Legacy alias: voiceId. |
lang |
SSML language code (default en-US). |
outputFormat |
Audio-file output format (default audio-24khz-48kbitrate-mono-mp3). |
voiceNoteOutputFormat |
Voice-note output format (default ogg-24khz-16bit-mono-opus). |
timeoutMs |
Request timeout override in milliseconds. Falls back to the global tts.timeoutMs. |
The provider is considered configured once apiKey is set plus one of
region, endpoint, or baseUrl. Env vars are only checked as a fallback
for config keys left unset. Workspace .env files cannot set
AZURE_SPEECH_ENDPOINT; use the process environment, global runtime dotenv,
or explicit config for endpoint routing.
참고
인증
Azure Speech uses a Speech resource key, not an Azure OpenAI key. The key
is sent as `Ocp-Apim-Subscription-Key`; OpenClaw derives
`https://<region>.tts.speech.microsoft.com` from `region` unless you
provide `endpoint` or `baseUrl`.
Voice names
Use the Azure Speech voice `ShortName` value, for example
`en-US-JennyNeural`. The bundled provider can list voices through the
same Speech resource and filters out voices marked deprecated, retired,
or disabled.
Audio outputs
Azure accepts output formats such as `audio-24khz-48kbitrate-mono-mp3`,
`ogg-24khz-16bit-mono-opus`, and `riff-24khz-16bit-mono-pcm`. OpenClaw
requests Ogg/Opus for `voice-note` targets so channels can send native
voice bubbles without an extra MP3 conversion, and forces
`raw-8khz-8bit-mono-mulaw` for telephony targets.
Alias
`azure` is accepted as a provider alias for existing config, but new
config should use `azure-speech` to avoid confusion with Azure OpenAI
model providers.
관련 문서
TTS overview, providers, and `tts` config.
Full config reference including `tts` settings.
All bundled OpenClaw providers.
Common issues and debugging steps.
검증 체크리스트
-
openclaw models list --provider azure-speech로 모델이 보이는지 확인 - 기본 model alias를
agents.defaults.model.primary에 설정했는지 확인 - API key / OAuth / 로컬 런타임 인증 경로를 공식 문서와 대조했는지 확인
- failover 시 데이터가 다른 provider로 이동할 수 있는지 검토