Talking with a robot
Summary
Demonstrates hybrid AI systems combining speech, vision, and language models to control a physical robot arm. Integrates speech recognition (whisper), speech synthesis (piper), large language model (llama3.1), and computer vision (yolo) into a unified system. All models run locally on consumer-grade hardware for privacy.
Themes
Keywords
multimodal AI, robot control, speech recognition, LLM, computer vision
Poster
Click image to open full size