Directory → SKILL
SKILLoptionalHermes Optional Skills

Llava

Vision-language chat: VQA, captioning, image dialogue

LLaVAVision-LanguageMultimodalVisual Question AnsweringImage ChatCLIPVicunaConversational AI
Last registry verification2026-08-18v1.0.0Orchestra Research
Plain meaning

What does it add to Hermes?

Vision-language chat: VQA, captioning, image dialogue

Llava is a skill related to design and media. It adds tools to create, read, or modify media such as designs, images, audio, or video.

This plain-language explanation is based on the publisher description. The original text remains visible for verification.

Use it when

Use it when your goal in design and media is clear and you can limit it to the data and actions it actually needs.

Skip it when

Do not add it merely to experiment when Hermes already has a simpler path, or when you cannot review its source and permissions.

Who is it for?

Best for users who want a repeatable way of working inside Hermes.

Safe first test

Start with a disposable asset and create a copy instead of changing the original.

Original publisher description

Vision-language chat: VQA, captioning, image dialogue

✓
Data source

This entry was indexed from Hermes Optional Skills. Our explanation interprets the type and domain without inventing a capability not present upstream.

!
Security review

The source is official or editorially reviewed, but you still need to review permissions and version compatibility.

Safe setup path

Inspect, install, then test.

  1. 01
    Open the source

    Match the publisher, license, and description to your need. Check the real update history.

  2. 02
    Review permissions and secrets

    Never paste a secret value into this site. Use environment-variable names and grant the smallest scope.

  3. 03
    Copy setup only after review

    The controls below copy text. They do not execute commands on your device.

  4. 04
    Test with a non-sensitive task

    Inspect the visible tools, then exclude write or delete tools you do not need.

Install command

Review the command, then copy it.

hermes skills install llava

Hermes Belarabi does not execute this command. Installation happens on your device and remains subject to Hermes scanning and your review.

The full skill definition

Exactly what Hermes loads when this skill runs.

Reproduced from the official documentation. Read it before enabling the skill: this text becomes the agent's instructions.

Vision-language chat: VQA, captioning, image dialogue.

Skill metadata

A lookup table. Do not read it all; find the row that applies to you.

SourceOptional — install with hermes skills install official/mlops/llava
Pathoptional-skills/mlops/llava
Version1.0.0
AuthorOrchestra Research
LicenseMIT
Dependenciestransformers, torch, pillow
Platformslinux, macos, windows
TagsLLaVA, Vision-Language, Multimodal, Visual Question Answering, Image Chat, CLIP, Vicuna, Conversational AI, Instruction Tuning, VQA

Reference: full SKILL.md

Explains the idea itself. Read it slowly; the later sections build on it.

Open-source vision-language model for conversational image understanding.

When to use LLaVA

Explains the idea itself. Read it slowly; the later sections build on it.

Use when:

  • Building vision-language chatbots
  • Visual question answering (VQA)
  • Image description and captioning
  • Multi-turn image conversations
  • Visual instruction following
  • Document understanding with images

Metrics:

  • 23,000+ GitHub stars
  • GPT-4V level capabilities (targeted)
  • Apache 2.0 License
  • Multiple model sizes (7B-34B params)

Use alternatives instead:

  • GPT-4V: Highest quality, API-based
  • CLIP: Simple zero-shot classification
  • BLIP-2: Better for captioning only
  • Flamingo: Research, not open-source

Quick start

Ordered, practical steps. Run one and confirm it worked before moving on.

Installation

Shell6 lines
# Clone repository
git clone https://github.com/haotian-liu/LLaVA
cd LLaVA

# Install
pip install -e .

Basic usage

Python40 lines
from llava.model.builder import load_pretrained_model
from llava.mm_utils import get_model_name_from_path, process_images, tokenizer_image_token
from llava.constants import IMAGE_TOKEN_INDEX, DEFAULT_IMAGE_TOKEN
from llava.conversation import conv_templates
from PIL import Image


# Load model
model_path = "liuhaotian/llava-v1.5-7b"
tokenizer, model, image_processor, context_len = load_pretrained_model(
    model_path=model_path,
    model_base=None,
    model_name=get_model_name_from_path(model_path)
)

# Load image
image = Image.open("image.jpg")
image_tensor = process_images([image], image_processor, model.config)
image_tensor = image_tensor.to(model.device, dtype=torch.float16)

# Create conversation
conv = conv_templates["llava_v1"].copy()
conv.append_message(conv.roles[0], DEFAULT_IMAGE_TOKEN + "\nWhat is in this image?")
conv.append_message(conv.roles[1], None)
prompt = conv.get_prompt()

# Generate response
input_ids = tokenizer_image_token(prompt, tokenizer, IMAGE_TOKEN_INDEX, return_tensors='pt').unsqueeze(0).to(model.device)

with torch.inference_mode():
    output_ids = model.generate(
        input_ids,
        images=image_tensor,
        do_sample=True,
        temperature=0.2,
        max_new_tokens=512
    )

response = tokenizer.decode(output_ids[0], skip_special_tokens=True).strip()
print(response)

Available models

Explains the idea itself. Read it slowly; the later sections build on it.

ModelParametersVRAMQuality
LLaVA-v1.5-7B7B~14 GBGood
LLaVA-v1.5-13B13B~28 GBBetter
LLaVA-v1.6-34B34B~70 GBBest
Python7 lines
# Load different models
model_7b = "liuhaotian/llava-v1.5-7b"
model_13b = "liuhaotian/llava-v1.5-13b"
model_34b = "liuhaotian/llava-v1.6-34b"

# 4-bit quantization for lower VRAM
load_4bit = True  # Reduces VRAM by ~4×

CLI usage

Explains the idea itself. Read it slowly; the later sections build on it.

Shell11 lines
# Single image query
python -m llava.serve.cli \
    --model-path liuhaotian/llava-v1.5-7b \
    --image-file image.jpg \
    --query "What is in this image?"

# Multi-turn conversation
python -m llava.serve.cli \
    --model-path liuhaotian/llava-v1.5-7b \
    --image-file image.jpg
# Then type questions interactively

Web UI (Gradio)

Explains the idea itself. Read it slowly; the later sections build on it.

Shell6 lines
# Launch Gradio interface
python -m llava.serve.gradio_web_server \
    --model-path liuhaotian/llava-v1.5-7b \
    --load-4bit  # Optional: reduce VRAM

# Access at http://localhost:7860

Multi-turn conversations

Settings you configure once. Change one at a time so you can see what each does. Set DEFAULT_IMAGE_TOKEN in your environment, not in the chat.

Python19 lines
# Initialize conversation
conv = conv_templates["llava_v1"].copy()

# Turn 1
conv.append_message(conv.roles[0], DEFAULT_IMAGE_TOKEN + "\nWhat is in this image?")
conv.append_message(conv.roles[1], None)
response1 = generate(conv, model, image)  # "A dog playing in a park"

# Turn 2
conv.messages[-1][1] = response1  # Add previous response
conv.append_message(conv.roles[0], "What breed is the dog?")
conv.append_message(conv.roles[1], None)
response2 = generate(conv, model, image)  # "Golden Retriever"

# Turn 3
conv.messages[-1][1] = response2
conv.append_message(conv.roles[0], "What time of day is it?")
conv.append_message(conv.roles[1], None)
response3 = generate(conv, model, image)

Common tasks

Explains the idea itself. Read it slowly; the later sections build on it.

Image captioning

Python2 lines
question = "Describe this image in detail."
response = ask(model, image, question)

Visual question answering

Python2 lines
question = "How many people are in the image?"
response = ask(model, image, question)

Object detection (textual)

Python2 lines
question = "List all the objects you can see in this image."
response = ask(model, image, question)

Scene understanding

Python2 lines
question = "What is happening in this scene?"
response = ask(model, image, question)

Document understanding

Python2 lines
question = "What is the main topic of this document?"
response = ask(model, document_image, question)

Training custom model

Explains the idea itself. Read it slowly; the later sections build on it.

Shell5 lines
# Stage 1: Feature alignment (558K image-caption pairs)
bash scripts/v1_5/pretrain.sh

# Stage 2: Visual instruction tuning (150K instruction data)
bash scripts/v1_5/finetune.sh

Quantization (reduce VRAM)

Explains the idea itself. Read it slowly; the later sections build on it.

Python10 lines
# 4-bit quantization
tokenizer, model, image_processor, context_len = load_pretrained_model(
    model_path="liuhaotian/llava-v1.5-13b",
    model_base=None,
    model_name=get_model_name_from_path("liuhaotian/llava-v1.5-13b"),
    load_4bit=True  # Reduces VRAM ~4×
)

# 8-bit quantization
load_8bit=True  # Reduces VRAM ~2×

Best practices

Explains the idea itself. Read it slowly; the later sections build on it.

  1. Start with 7B model - Good quality, manageable VRAM
  2. Use 4-bit quantization - Reduces VRAM significantly
  3. GPU required - CPU inference extremely slow
  4. Clear prompts - Specific questions get better answers
  5. Multi-turn conversations - Maintain conversation context
  6. Temperature 0.2-0.7 - Balance creativity/consistency
  7. max_new_tokens 512-1024 - For detailed responses
  8. Batch processing - Process multiple images sequentially

Performance

Explains the idea itself. Read it slowly; the later sections build on it.

ModelVRAM (FP16)VRAM (4-bit)Speed (tokens/s)
7B~14 GB~4 GB~20
13B~28 GB~8 GB~12
34B~70 GB~18 GB~5

On A100 GPU

Benchmarks

Explains the idea itself. Read it slowly; the later sections build on it.

LLaVA achieves competitive scores on:

  • VQAv2: 78.5%
  • GQA: 62.0%
  • MM-Vet: 35.4%
  • MMBench: 64.3%

Limitations

Explains the idea itself. Read it slowly; the later sections build on it.

  1. Hallucinations - May describe things not in image
  2. Spatial reasoning - Struggles with precise locations
  3. Small text - Difficulty reading fine print
  4. Object counting - Imprecise for many objects
  5. VRAM requirements - Need powerful GPU
  6. Inference speed - Slower than CLIP

Integration with frameworks

Explains the idea itself. Read it slowly; the later sections build on it.

LangChain

Python8 lines
from langchain.llms.base import LLM

class LLaVALLM(LLM):
    def _call(self, prompt, stop=None):
        # Custom LLaVA inference
        return response

llm = LLaVALLM()

Gradio App

Python12 lines


def chat(image, text, history):
    response = ask_llava(model, image, text)
    return response

demo = gr.ChatInterface(
    chat,
    additional_inputs=[gr.Image(type="pil")],
    title="LLaVA Chat"
)
demo.launch()

Resources

Explains the idea itself. Read it slowly; the later sections build on it.

  • GitHub: https://github.com/haotian-liu/LLaVA ⭐ 23,000+
  • Paper: https://arxiv.org/abs/2304.08485
  • Demo: https://llava.hliu.cc
  • Models: https://huggingface.co/liuhaotian
  • License: Apache 2.0