Academy → Practical GuidesOfficial documentation · Arabic guidance

Google Vertex AI

التشغيل عبر Google Vertex AI

Intermediate5 min readLesson 234 questions✓ 2026-08-18
Before you read

What this page is, and what it holds.

This page covers Google Vertex AI. You will use hermes model and hermes setup here; about 5 minutes to read. The priciest model is not always best for your task. Compare on one task and set a spend cap.

9sections
6code examples
1tables
4commands
794source words
The official one-line description

Use Hermes Agent with Gemini on Google Cloud Vertex AI — OAuth2 service account or ADC, GCP billing and quotas, no static API key

What you will be able to do

Outcomes taken from this page, not a template.

  • Understand what المزوّد والنموذج is and when you need it.
  • Run hermes model and hermes setup and understand what happens next.
  • Read the table and take only the row that applies to you.
  • Set VERTEX_CREDENTIALS_PATH in the right place.
Identifiers you will meet

Exactly as they appear in Hermes.

Commands
  • hermes model
  • hermes setup
  • hermes chat
  • hermes doctor
Environment variables
  • VERTEX_CREDENTIALS_PATH
  • GOOGLE_APPLICATION_CREDENTIALS
  • VERTEX_PROJECT_ID
  • VERTEX_REGION
Page map

Jump to the part you need.

  1. 01Prerequisites
  2. 02Quick Start
  3. 03Configuration
  4. 04Available Models
  5. 05Switching Models Mid-Session
  6. 06Reasoning / Thinking
  7. 07Diagnostics
  8. 08Troubleshooting
  9. 09Related
The full official page

Nothing summarised away.

The documentation body below is reproduced from the official source so commands and identifiers stay exact. Each section carries a short note describing what it contains.

Hermes Agent supports Gemini models on Google Cloud Vertex AI through Vertex's OpenAI-compatible endpoint. Unlike the Google AI Studio provider (which uses a static API key against generativelanguage.googleapis.com), Vertex gives you enterprise-grade rate limits and GCP billing/credits, and is the right choice when you want Gemini usage to draw on your Google Cloud account rather than an AI Studio key.

Prerequisites

Commands you type in a terminal. Understand what one does before copying it. Commands here: hermes setup.

  • A Google Cloud project with the Vertex AI API enabled and billing active.
  • Credentials, one of:
  • a service-account JSON key file with the roles/aiplatform.user role, or
  • Application Default Credentials via gcloud auth application-default login (or the metadata server when running on a GCP VM).
  • google-auth — installed automatically the first time you select Vertex (lazy install). Run hermes setup to repair a managed install if that fails.

Quick Start

Ordered, practical steps. Run one and confirm it worked before moving on. Commands here: hermes chat, hermes model.

Shell15 lines
# Option A — service account JSON (recommended for servers / gateways)
echo "VERTEX_CREDENTIALS_PATH=/path/to/service-account.json" >> ~/.hermes/.env

# Option B — Application Default Credentials (good for local dev)
gcloud auth application-default login

# Select Vertex as your provider
hermes model
# → Choose "More providers..." → "Google Vertex AI"
# → Enter your GCP project ID (or leave blank to use the one in your credentials)
# → Choose a region (default: global)
# → Select a Gemini model

# Start chatting
hermes chat

Configuration

Settings you configure once. Change one at a time so you can see what each does. Set VERTEX_CREDENTIALS_PATH, GOOGLE_APPLICATION_CREDENTIALS in your environment, not in the chat.

Vertex splits its settings by sensitivity:

  • The credential path is a pointer to a secret and lives in ~/.hermes/.env.
  • Project ID and region are non-secret routing settings and live in ~/.hermes/config.yaml.

~/.hermes/.env:

Shell3 lines
# One of these (checked in this order); omit both to use ADC:
VERTEX_CREDENTIALS_PATH=/path/to/service-account.json
GOOGLE_APPLICATION_CREDENTIALS=/path/to/service-account.json

~/.hermes/config.yaml:

YAML7 lines
model:
  default: google/gemini-3-flash-preview
  provider: vertex

vertex:
  project_id: my-gcp-project   # blank → use the project embedded in the credentials
  region: global               # "global" is required for the Gemini 3.x previews

How authentication works

  1. Hermes resolves credentials in this order: VERTEX_CREDENTIALS_PATH → GOOGLE_APPLICATION_CREDENTIALS → ADC.
  2. It mints an OAuth2 access token (cloud-platform scope) and caches it, refreshing when the token is within 5 minutes of expiry.
  3. The token is handed to a standard OpenAI client pointed at the Vertex endpoint:
Text1 line
   https://aiplatform.googleapis.com/v1beta1/projects/{project}/locations/{region}/endpoints/openapi

Regional locations use a {region}-aiplatform.googleapis.com host instead.

  1. If a session runs longer than the token lifetime and a request returns 401, Hermes re-mints the token and retries automatically. On a long-running gateway, if ADC's refresh token has itself expired, Hermes falls back to the service-account JSON when one is configured.

Available Models

A lookup table. Do not read it all; find the row that applies to you. Commands here: hermes model.

Vertex requires the google/ vendor prefix on model IDs. The hermes model picker offers:

ModelID
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview
Gemini 3 Pro Previewgoogle/gemini-3-pro-preview
Gemini 3 Flash Previewgoogle/gemini-3-flash-preview
Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview
Gemini 2.5 Progoogle/gemini-2.5-pro
Gemini 2.5 Flashgoogle/gemini-2.5-flash

Switching Models Mid-Session

Commands you type in a terminal. Understand what one does before copying it. Commands here: hermes model.

Text2 lines
/model google/gemini-3-pro-preview
/model google/gemini-3-flash-preview

/model switches among already-configured providers and models; it does not collect new credentials. Configure Vertex with hermes model first.

Reasoning / Thinking

Explains the idea itself. Read it slowly; the later sections build on it.

Vertex exposes Gemini's thinking budget through the OpenAI-compatible surface. Hermes maps its reasoning-effort setting onto extra_body.google.thinking_config automatically, so reasoning_effort works the same way it does on other Gemini surfaces.

Diagnostics

Commands you type in a terminal. Understand what one does before copying it. Commands here: hermes doctor.

Shell1 line
hermes doctor

The doctor reports whether Vertex credentials can be resolved (service-account path or ADC) and whether the provider is configured.

Troubleshooting

A troubleshooting section. Find the symptom that matches yours rather than reading it end to end. Commands here: hermes setup.

"Vertex AI credentials could not be resolved"

Hermes found neither a service-account JSON nor working ADC. Either set VERTEX_CREDENTIALS_PATH in ~/.hermes/.env, or run gcloud auth application-default login. If your project isn't embedded in the credentials, set vertex.project_id in config.yaml.

google-auth not installed

Hermes lazy-installs it the first time you select the Vertex provider. If that fails, run hermes setup to repair the managed install.

404 on Gemini 3.x models

You are probably on a regional endpoint. Set region: global in the vertex: section of config.yaml (or unset VERTEX_REGION).

403 / permission denied

The service account (or your ADC identity) needs the roles/aiplatform.user role on the project, and the Vertex AI API must be enabled for that project.

Knowledge check

4 questions answered by this page alone.

Every option is a real identifier from the Hermes documentation. The wrong ones are real too, just from other pages.

1. In this lesson's table, what is the “ID” for “Gemini 2.5 Flash”?
2. Which of these environment variables actually appears in this lesson?
3. Which of these headings does not appear in this lesson?
4. Which configuration key appears in this lesson's examples?