Inference serves 4 of the models we track. We do not support bring-your-own-key for this provider yet.
| Model | Made by | Input / 1M | Output / 1M | Context |
|---|---|---|---|---|
| Llama 3.2 1b Instruct | NVIDIA | $0.01 | $0.01 | 16K |
| Llama 3.2 3B Instruct | NVIDIA | $0.02 | $0.02 | 16K |
| Llama 3.1 8B Instruct | NVIDIA | $0.025 | $0.025 | 16K |
| Llama 3.2 11b Vision Instruct | NVIDIA | $0.055 | $0.055 | 16K |
Inference serves 4 of the models we track, across 1 model type.
Llama 3.2 1b Instruct at $0.01 per million input tokens.
Not yet — Inference is not currently one of the providers Serenities supports for bring-your-own-key.