NVIDIA NIM replaced its depleting trial credits with a recurring per-account rate limit (40 RPM default, varies by model), verified June 2026. The trial ToS still scopes usage to evaluation/prototyping, not production.
OpenRouter’s :free daily cap (50/day, or 1000/day once you have ever bought $10 of credits) is shared across ALL :free models on the account, not per model. Per-row rpd values here are therefore optimistic; the router’s cooldown handling absorbs the shared 429s.
HuggingFace Inference Providers grants only ~$0.10/month of routed credit on the free tier (PRO is $2/month). Enough for light experimentation; exhausts quickly. Credits apply only to HF-routed requests.
FreeLLMAPI is a self-hosted router you run yourself. Install it, paste in your free NVIDIA NIM key, and google/gemma-4-31b-it answers on an OpenAI-compatible endpoint at http://localhost:3001/v1. No credit card, no hosted middleman: your prompts and your provider keys never leave your machine.
curl http://localhost:3001/v1/chat/completions \
-H "Authorization: Bearer freellmapi-your-unified-key" \
-H "Content-Type: application/json" \
-d '{
"model": "google/gemma-4-31b-it",
"messages": [{"role": "user", "content": "Say hi in five words."}]
}'
The same request in Python (openai)
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:3001/v1",
api_key="freellmapi-your-unified-key",
)
resp = client.chat.completions.create(
model="google/gemma-4-31b-it",
messages=[{"role": "user", "content": "Say hi in five words."}],
)
print(resp.choices[0].message.content)
The router answers on /v1/chat/completions and every other OpenAI surface, plus the Anthropic Messages API, so existing clients need only a new base_url. Swap the model id for auto and the router picks the best free model that is still under its limits. Full reference: docs/api.md.
Related free models
Same key, same router, no extra setup. Any of these is one model id away.