NVIDIA Provider
기준일: 2026-07-26
난이도: 중급
공식 기준: NVIDIA
개요
이 페이지는 OpenClaw NVIDIA provider의 인증, 모델 ref, 온보딩 플래그, 설정 키를 공식 문서 기준으로 정리합니다.
공식 요약: Use NVIDIA's OpenAI-compatible API in OpenClaw
model ref는 보통 provider/model 형식입니다. API key·endpoint·model id는 아래 원문 값을 그대로 쓰세요. 임의로 키나 엔드포인트를 만들지 마세요.
빠른 참조
| Model ref | Name | Context | Max output |
|---|---|---|---|
nvidia/nvidia/nemotron-3-ultra-550b-a55b |
Nemotron 3 Ultra 550B | 1,048,576 | 8,192 |
nvidia/nvidia/nemotron-3-super-120b-a12b |
Nemotron 3 Super 120B | 1,000,000 | 8,192 |
nvidia/z-ai/glm-5.2 |
GLM 5.2 | 202,752 | 8,192 |
nvidia/moonshotai/kimi-k2.6 |
Kimi K2.6 | 262,144 | 8,192 |
nvidia/minimaxai/minimax-m3 |
Minimax M3 | 196,608 | 8,192 |
nvidia/deepseek-ai/deepseek-v4-pro |
DeepSeek V4 Pro | 262,144 | 16,384 |
nvidia/qwen/qwen3.5-397b-a17b |
Qwen3.5 397B A17B | 262,144 | 16,384 |
공식 문서 기반 상세
아래는 공식 providers/nvidia 문서를 Mintlify 컴포넌트만 정리하고 공통 제목을 한국어로 맞춘 내용입니다. 설정 키, env, model ref, CLI 플래그는 원문 그대로 유지합니다.
NVIDIA serves open models for free through an OpenAI-compatible API at
https://integrate.api.nvidia.com/v1, authenticated with an API key from
build.nvidia.com. OpenClaw
defaults the NVIDIA provider to Nemotron 3 Ultra, NVIDIA's 550B total / 55B
active reasoning model for long-context agentic work.
시작하기
Get your API key
Create an API key at [build.nvidia.com](https://build.nvidia.com/settings/api-keys).
Export the key and run onboarding
```bash
export NVIDIA_API_KEY="nvapi-..."
openclaw onboard --auth-choice nvidia-api-key
```
Set an NVIDIA model
```bash
openclaw models set nvidia/nvidia/nemotron-3-ultra-550b-a55b
```
For non-interactive setup, pass the key directly:
openclaw onboard --auth-choice nvidia-api-key --nvidia-api-key "nvapi-..."
--nvidia-api-keylands the key in shell history andpsoutput. Prefer theNVIDIA_API_KEYenvironment variable when possible.
설정 예시
{
env: { NVIDIA_API_KEY: "nvapi-..." },
models: {
providers: {
nvidia: {
baseUrl: "https://integrate.api.nvidia.com/v1",
api: "openai-completions",
},
},
},
agents: {
defaults: {
model: { primary: "nvidia/nvidia/nemotron-3-ultra-550b-a55b" },
},
},
}
Featured catalog
When an NVIDIA API key is configured, setup and model-selection paths fetch
NVIDIA's public featured-model catalog from
https://assets.ngc.nvidia.com/products/api-catalog/featured-models.json and
cache the result for 24 hours (first 32 entries, imported as free text-input
rows). New featured models from build.nvidia.com therefore appear in setup and
model-selection surfaces without waiting for an OpenClaw release. When the
live feed is available, the first returned model is the preselected option
during NVIDIA setup.
The fetch uses a fixed HTTPS host policy for assets.ngc.nvidia.com. If no
NVIDIA API key is configured, or if the feed is unavailable or malformed,
OpenClaw falls back to the bundled catalog and bundled default below.
Nemotron 3 Ultra
Nemotron 3 Ultra is the default NVIDIA model in OpenClaw. NVIDIA's build page for
nvidia/nemotron-3-ultra-550b-a55b
lists it as an available free endpoint with a 1M-token context specification.
The bundled Ultra row sends
chat_template_kwargs: { enable_thinking: false, force_nonempty_content: true }
by default so normal chat output stays in the visible answer instead of
exposing reasoning text.
Use Ultra for the highest-capability NVIDIA default. Keep Super selected when you want the smaller Nemotron 3 option, or choose one of the third-party models hosted in NVIDIA's catalog when their context, latency, or behavior fits better.
Bundled fallback catalog
The selectable bundled rows snapshot NVIDIA's featured-model catalog. Deprecated compatibility rows remain resolvable by exact reference but stay out of model pickers.
| Model ref | Name | Context | Max output |
|---|---|---|---|
nvidia/nvidia/nemotron-3-ultra-550b-a55b |
Nemotron 3 Ultra 550B | 1,048,576 | 8,192 |
nvidia/nvidia/nemotron-3-super-120b-a12b |
Nemotron 3 Super 120B | 1,000,000 | 8,192 |
nvidia/z-ai/glm-5.2 |
GLM 5.2 | 202,752 | 8,192 |
nvidia/moonshotai/kimi-k2.6 |
Kimi K2.6 | 262,144 | 8,192 |
nvidia/minimaxai/minimax-m3 |
Minimax M3 | 196,608 | 8,192 |
nvidia/deepseek-ai/deepseek-v4-pro |
DeepSeek V4 Pro | 262,144 | 16,384 |
nvidia/qwen/qwen3.5-397b-a17b |
Qwen3.5 397B A17B | 262,144 | 16,384 |
The full compatibility catalog also retains these shipped refs for existing
configurations: nvidia/moonshotai/kimi-k2.5, nvidia/z-ai/glm-5.1,
nvidia/minimaxai/minimax-m2.5, nvidia/z-ai/glm5, and
nvidia/minimaxai/minimax-m2.7. They remain available by exact reference but
never appear in onboarding or model pickers.
고급 설정
Auto-enable behavior
The provider auto-enables when the `NVIDIA_API_KEY` environment variable is
set or a key was stored during onboarding. No explicit provider config is
required beyond the key.
Catalog and pricing
OpenClaw prefers NVIDIA's public featured-model catalog when NVIDIA auth is
configured and caches it for 24 hours. The bundled selectable fallback is a
static snapshot of NVIDIA's featured-model catalog; deprecated exact-reference
compatibility rows are hidden from model pickers. Costs default to `0` in
source since NVIDIA currently offers free API access for the listed models.
OpenAI-compatible endpoint
OpenClaw talks to NVIDIA with the `openai-completions` adapter against the
standard `/v1` chat completions route. Any OpenAI-compatible tooling should
work out of the box with the NVIDIA base URL.
Nemotron 3 Ultra reasoning params
NVIDIA's Ultra sample request uses `chat_template_kwargs.enable_thinking`
and `reasoning_budget` for reasoning output. OpenClaw's bundled Ultra row
disables template thinking by default for normal chat use. If you need to
opt into NVIDIA reasoning output or force other NVIDIA-specific request
fields, set per-model params and keep provider-specific overrides scoped to
the NVIDIA model:
```json5
{
agents: {
defaults: {
models: {
"nvidia/nvidia/nemotron-3-ultra-550b-a55b": {
params: {
chat_template_kwargs: { enable_thinking: true },
extra_body: { reasoning_budget: 16384 },
},
},
},
},
},
}
```
`params.chat_template_kwargs` merges into any `chat_template_kwargs`
already on the request instead of replacing the whole object.
`params.extra_body` is the final OpenAI-compatible request-body override
and overwrites colliding payload keys, so use it only for fields NVIDIA
documents for the selected endpoint.
Slow custom provider responses
Some NVIDIA-hosted custom models can take longer than the default ~120s
model idle watchdog before they emit a first response chunk. For custom
NVIDIA provider entries, raise the provider timeout instead of the whole
agent runtime timeout; `timeoutSeconds` covers provider HTTP requests and
raises the idle/stream watchdog ceiling for that provider:
```json5
{
models: {
providers: {
"custom-integrate-api-nvidia-com": {
baseUrl: "https://integrate.api.nvidia.com/v1",
api: "openai-completions",
apiKey: "NVIDIA_API_KEY",
timeoutSeconds: 300,
},
},
},
agents: {
defaults: {
models: {
"custom-integrate-api-nvidia-com/meta/llama-3.1-70b-instruct": {
params: { thinking: "off" },
},
},
},
},
}
```
NVIDIA models are currently free to use. Check build.nvidia.com for the latest availability and rate-limit details.
관련 문서
Choosing providers, model refs, and failover behavior.
Full config reference for agents, models, and providers.
검증 체크리스트
-
openclaw models list --provider nvidia로 모델이 보이는지 확인 - 기본 model alias를
agents.defaults.model.primary에 설정했는지 확인 - API key / OAuth / 로컬 런타임 인증 경로를 공식 문서와 대조했는지 확인
- failover 시 데이터가 다른 provider로 이동할 수 있는지 검토