Skip to content

Voice Control Tutorial

Learn how to use voice commands for hands-free interaction with CatGo.

Overview

CatGo supports voice input via Whisper (local speech-to-text) and voice output via text-to-speech, enabling conversational interaction.

Step 1: Enable Voice

Open the Gesture Settings pane and enable voice input. Select your microphone.

Step 2: Speech-to-Text Setup

Whisper Engine

CatGo uses a local Whisper model for privacy-preserving speech recognition. The model is downloaded on first use.

Step 3: Voice Commands

Structure Commands

  • "Rotate left/right/up/down"
  • "Zoom in/out"
  • "Reset view"
  • "Show bonds/labels/axes"

Atom Art

  • "Place a carbon atom"
  • "Add oxygen here"
  • "Build a benzene ring"

Analysis Commands

  • "Compute RDF"
  • "Show band structure"
  • "Optimize structure"

Step 4: Text-to-Speech Response

CatGo responds with voice feedback for executed commands.

Step 5: Customize

Adjust voice activation sensitivity, language, and TTS voice in settings.

CatGo is licensed under AGPL-3.0-or-later. If CatGo contributes to your work, please include “This work used CatGo (https://app.catgo-ucsd.org).” and the preferred citation in CITATION.cff. This request is not an additional condition of the AGPL license.