Google Vertex AI
التشغيل عبر Google Vertex AI
What this page is, and what it holds.
This page covers Google Vertex AI. You will use hermes model and hermes setup here; about 5 minutes to read. The priciest model is not always best for your task. Compare on one task and set a spend cap.
Use Hermes Agent with Gemini on Google Cloud Vertex AI — OAuth2 service account or ADC, GCP billing and quotas, no static API key
Outcomes taken from this page, not a template.
- Understand what المزوّد والنموذج is and when you need it.
- Run
hermes modelandhermes setupand understand what happens next. - Read the table and take only the row that applies to you.
- Set
VERTEX_CREDENTIALS_PATHin the right place.
Exactly as they appear in Hermes.
hermes modelhermes setuphermes chathermes doctor
VERTEX_CREDENTIALS_PATHGOOGLE_APPLICATION_CREDENTIALSVERTEX_PROJECT_IDVERTEX_REGION
Jump to the part you need.
Nothing summarised away.
The documentation body below is reproduced from the official source so commands and identifiers stay exact. Each section carries a short note describing what it contains.
Hermes Agent supports Gemini models on Google Cloud Vertex AI through Vertex's OpenAI-compatible endpoint. Unlike the Google AI Studio provider (which uses a static API key against generativelanguage.googleapis.com), Vertex gives you enterprise-grade rate limits and GCP billing/credits, and is the right choice when you want Gemini usage to draw on your Google Cloud account rather than an AI Studio key.
Prerequisites
Commands you type in a terminal. Understand what one does before copying it. Commands here: hermes setup.
- A Google Cloud project with the Vertex AI API enabled and billing active.
- Credentials, one of:
- a service-account JSON key file with the
roles/aiplatform.userrole, or - Application Default Credentials via
gcloud auth application-default login(or the metadata server when running on a GCP VM). google-auth— installed automatically the first time you select Vertex (lazy install). Runhermes setupto repair a managed install if that fails.
Quick Start
Ordered, practical steps. Run one and confirm it worked before moving on. Commands here: hermes chat, hermes model.
# Option A — service account JSON (recommended for servers / gateways)
echo "VERTEX_CREDENTIALS_PATH=/path/to/service-account.json" >> ~/.hermes/.env
# Option B — Application Default Credentials (good for local dev)
gcloud auth application-default login
# Select Vertex as your provider
hermes model
# → Choose "More providers..." → "Google Vertex AI"
# → Enter your GCP project ID (or leave blank to use the one in your credentials)
# → Choose a region (default: global)
# → Select a Gemini model
# Start chatting
hermes chatConfiguration
Settings you configure once. Change one at a time so you can see what each does. Set VERTEX_CREDENTIALS_PATH, GOOGLE_APPLICATION_CREDENTIALS in your environment, not in the chat.
Vertex splits its settings by sensitivity:
- The credential path is a pointer to a secret and lives in
~/.hermes/.env. - Project ID and region are non-secret routing settings and live in
~/.hermes/config.yaml.
~/.hermes/.env:
# One of these (checked in this order); omit both to use ADC:
VERTEX_CREDENTIALS_PATH=/path/to/service-account.json
GOOGLE_APPLICATION_CREDENTIALS=/path/to/service-account.json~/.hermes/config.yaml:
model:
default: google/gemini-3-flash-preview
provider: vertex
vertex:
project_id: my-gcp-project # blank → use the project embedded in the credentials
region: global # "global" is required for the Gemini 3.x previewsHow authentication works
- Hermes resolves credentials in this order:
VERTEX_CREDENTIALS_PATH→GOOGLE_APPLICATION_CREDENTIALS→ ADC. - It mints an OAuth2 access token (
cloud-platformscope) and caches it, refreshing when the token is within 5 minutes of expiry. - The token is handed to a standard OpenAI client pointed at the Vertex endpoint:
https://aiplatform.googleapis.com/v1beta1/projects/{project}/locations/{region}/endpoints/openapiRegional locations use a {region}-aiplatform.googleapis.com host instead.
- If a session runs longer than the token lifetime and a request returns
401, Hermes re-mints the token and retries automatically. On a long-running gateway, if ADC's refresh token has itself expired, Hermes falls back to the service-account JSON when one is configured.
Available Models
A lookup table. Do not read it all; find the row that applies to you. Commands here: hermes model.
Vertex requires the google/ vendor prefix on model IDs. The hermes model picker offers:
| Model | ID |
|---|---|
| Gemini 3.1 Pro Preview | google/gemini-3.1-pro-preview |
| Gemini 3 Pro Preview | google/gemini-3-pro-preview |
| Gemini 3 Flash Preview | google/gemini-3-flash-preview |
| Gemini 3.1 Flash Lite Preview | google/gemini-3.1-flash-lite-preview |
| Gemini 2.5 Pro | google/gemini-2.5-pro |
| Gemini 2.5 Flash | google/gemini-2.5-flash |
Switching Models Mid-Session
Commands you type in a terminal. Understand what one does before copying it. Commands here: hermes model.
/model google/gemini-3-pro-preview
/model google/gemini-3-flash-preview/model switches among already-configured providers and models; it does not collect new credentials. Configure Vertex with hermes model first.
Reasoning / Thinking
Explains the idea itself. Read it slowly; the later sections build on it.
Vertex exposes Gemini's thinking budget through the OpenAI-compatible surface. Hermes maps its reasoning-effort setting onto extra_body.google.thinking_config automatically, so reasoning_effort works the same way it does on other Gemini surfaces.
Diagnostics
Commands you type in a terminal. Understand what one does before copying it. Commands here: hermes doctor.
hermes doctorThe doctor reports whether Vertex credentials can be resolved (service-account path or ADC) and whether the provider is configured.
Troubleshooting
A troubleshooting section. Find the symptom that matches yours rather than reading it end to end. Commands here: hermes setup.
"Vertex AI credentials could not be resolved"
Hermes found neither a service-account JSON nor working ADC. Either set VERTEX_CREDENTIALS_PATH in ~/.hermes/.env, or run gcloud auth application-default login. If your project isn't embedded in the credentials, set vertex.project_id in config.yaml.
google-auth not installed
Hermes lazy-installs it the first time you select the Vertex provider. If that fails, run hermes setup to repair the managed install.
404 on Gemini 3.x models
You are probably on a regional endpoint. Set region: global in the vertex: section of config.yaml (or unset VERTEX_REGION).
403 / permission denied
The service account (or your ADC identity) needs the roles/aiplatform.user role on the project, and the Vertex AI API must be enabled for that project.
4 questions answered by this page alone.
Every option is a real identifier from the Hermes documentation. The wrong ones are real too, just from other pages.