Use it when your goal in design and media is clear and you can limit it to the data and actions it actually needs.
Songsee
Audio spectrograms/features (mel, chroma, MFCC) via CLI
What does it add to Hermes?
Audio spectrograms/features (mel, chroma, MFCC) via CLI
Songsee is a skill related to design and media. It adds tools to create, read, or modify media such as designs, images, audio, or video.
This plain-language explanation is based on the publisher description. The original text remains visible for verification.
Do not add it merely to experiment when Hermes already has a simpler path, or when you cannot review its source and permissions.
Best for users who want a repeatable way of working inside Hermes.
Start with a disposable asset and create a copy instead of changing the original.
Audio spectrograms/features (mel, chroma, MFCC) via CLI
This entry was indexed from Hermes Bundled Skills. Our explanation interprets the type and domain without inventing a capability not present upstream.
The source is official or editorially reviewed, but you still need to review permissions and version compatibility.
Inspect, install, then test.
- 01Open the source
Match the publisher, license, and description to your need. Check the real update history.
- 02Review permissions and secrets
Never paste a secret value into this site. Use environment-variable names and grant the smallest scope.
- 03Copy setup only after review
The controls below copy text. They do not execute commands on your device.
- 04Test with a non-sensitive task
Inspect the visible tools, then exclude write or delete tools you do not need.
Already installed with Hermes.
This skill ships with Hermes and loads when the agent decides it is relevant. There is nothing to install; read the definition below so you know what it will do.
Open the official page ↗Exactly what Hermes loads when this skill runs.
Reproduced from the official documentation. Read it before enabling the skill: this text becomes the agent's instructions.
Audio spectrograms/features (mel, chroma, MFCC) via CLI.
Skill metadata
A lookup table. Do not read it all; find the row that applies to you.
| Source | Bundled (installed by default) |
| Path | skills/media/songsee |
| Version | 1.0.0 |
| Author | community |
| License | MIT |
| Platforms | linux, macos, windows |
| Tags | Audio, Visualization, Spectrogram, Music, Analysis |
Reference: full SKILL.md
Explains the idea itself. Read it slowly; the later sections build on it.
Generate spectrograms and multi-panel audio feature visualizations from audio files.
Prerequisites
Explains the idea itself. Read it slowly; the later sections build on it.
Requires Go ↗:
go install github.com/steipete/songsee/cmd/songsee@latestOptional: ffmpeg for formats beyond WAV/MP3.
Quick Start
Ordered, practical steps. Run one and confirm it worked before moving on.
# Basic spectrogram
songsee track.mp3
# Save to specific file
songsee track.mp3 -o spectrogram.png
# Multi-panel visualization grid
songsee track.mp3 --viz spectrogram,mel,chroma,hpss,selfsim,loudness,tempogram,mfcc,flux
# Time slice (start at 12.5s, 8s duration)
songsee track.mp3 --start 12.5 --duration 8 -o slice.jpg
# From stdin
cat track.mp3 | songsee - --format png -o out.pngVisualization Types
A lookup table. Do not read it all; find the row that applies to you.
Use --viz with comma-separated values:
| Type | Description |
|---|---|
spectrogram | Standard frequency spectrogram |
mel | Mel-scaled spectrogram |
chroma | Pitch class distribution |
hpss | Harmonic/percussive separation |
selfsim | Self-similarity matrix |
loudness | Loudness over time |
tempogram | Tempo estimation |
mfcc | Mel-frequency cepstral coefficients |
flux | Spectral flux (onset detection) |
Multiple --viz types render as a grid in a single image.
Common Flags
A lookup table. Do not read it all; find the row that applies to you.
| Flag | Description |
|---|---|
--viz | Visualization types (comma-separated) |
--style | Color palette: classic, magma, inferno, viridis, gray |
--width / --height | Output image dimensions |
--window / --hop | FFT window and hop size |
--min-freq / --max-freq | Frequency range filter |
--start / --duration | Time slice of the audio |
--format | Output format: jpg or png |
-o | Output file path |
Notes
Explains the idea itself. Read it slowly; the later sections build on it.
- WAV and MP3 are decoded natively; other formats require
ffmpeg - Output images can be inspected with
vision_analyzefor automated audio analysis - Useful for comparing audio outputs, debugging synthesis, or documenting audio processing pipelines