Projects · · 4 min read
How I built JARVIS, my own AI assistant in Python
I built JARVIS, a free‑tool Python voice assistant that can chat, listen, and search the web. Here’s how I did it step by step, using VS Code, Claude, Supabase and Gemini.
I wanted to build python voice assistant that could understand my spoken commands, reply in text or speech, and fetch answers from the web without spending a dime. The idea started when I realized I was juggling multiple tabs and notebooks, and a single assistant could keep everything in one place. This post walks through the exact steps I followed, the free services I chose, and the hiccups I still have to iron out.
Step-by-step guide to build python voice assistant
The overall plan was simple: a loop that records my voice, sends it to a speech‑to‑text model, forwards the transcript to a chat model, turns the response back into speech, and finally runs a web search if the answer looks incomplete. I kept the architecture linear so I could add or remove parts later. The first version only handled text chat; voice and search came in the next two iterations.
Free tools and APIs I used
All the components have a free tier, which matches my rule of staying cost‑free. I wrote the code in VS Code because the editor is lightweight and works well on my old laptop. Inside VS Code I installed the Claude extension, which helped me generate snippets and debug on the fly. For speech‑to‑text I used the free tier of Gemini’s whisper model, and for text‑to‑speech I relied on the built‑in macOS say command (no API key needed). Web search is powered by Supabase Edge Functions that call the DuckDuckGo Instant Answer API, both of which have generous free limits. I also read about similar projects in my earlier post about a memory‑match game [/blog/what-a-memory-match-game-taught-me-about-state] and about how I use AI tools without outsourcing my thinking [/blog/how-i-use-ai-tools-without-outsourcing-my-thinking]. You can see more of my experiments on my projects page [/projects].
- VS Code with Claude extension
- Gemini whisper (free tier) for speech‑to‑text
- macOS say command for text‑to‑speech
- Supabase Edge Functions + DuckDuckGo Instant Answer API for web search
Writing the chat and voice loops with Claude
I started by drafting a simple REPL that accepted text input and printed the response from Claude. When the loop worked, I asked Claude to wrap the input in a function that records audio, sends it to Gemini, and returns the transcript. Claude generated the code, I copied it into my file, and then I ran it. The first run failed because the audio library needed ffmpeg, so I installed it and tried again. Each error became a quick prompt to Claude: "Fix the ImportError for ffmpeg". Within a few hours I had a function called get_voice_input() that returned a clean string.
import subprocess, json, os
def get_voice_input():
# Record 5 seconds of audio to temp.wav
subprocess.run(['ffmpeg', '-y', '-f', 'avfoundation', '-i', ':0', '-t', '5', 'temp.wav'], stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
# Send to Gemini Whisper (pseudo code)
with open('temp.wav', 'rb') as f:
audio_bytes = f.read()
transcript = gemini_whisper_api(audio_bytes)
return transcript.strip()
Adding web search with Supabase and Gemini
Once the chat loop was stable, I noticed many of my questions needed up‑to‑date facts. I added a small check: if the response contains the phrase "I don’t know" or looks like a generic answer, I call a Supabase Edge Function that forwards the query to DuckDuckGo. The function returns a short snippet, which I then feed back to the chat model for a polished reply. This kept the user experience smooth and avoided hard‑coding any API keys in the client script.
Putting everything together in VS Code
With the three pieces – voice input, chat response, and optional web search – I stitched them into a single async main loop. I used Python’s asyncio library so the voice recording, API calls, and speech synthesis could happen without blocking each other. Running the script in the integrated terminal showed the assistant responding in real time. I added a few safety prints to see which path (chat only or chat + search) was taken, which helped me debug later when the search API throttled.
What I’m adding next
The current version works for simple queries, but I still want better context handling – the assistant should remember the last few interactions. I also plan to replace the macOS say command with a cross‑platform TTS service so I can run JARVIS on Linux. Finally, I want to expose a small web UI so I can trigger the assistant from my phone without opening a terminal.
Takeaway
Building JARVIS taught me that a free‑tool stack can get you far if you break the problem into tiny steps and let an AI helper fill the gaps. The hardest part was wiring the pieces together, not writing the individual functions. If you’re starting from scratch, focus on one capability at a time – text chat, then voice, then search – and keep the code modular. You’ll be surprised how quickly a functional assistant emerges.
If you want to keep reading, I also wrote about What a memory match game taught me about state and Learning to code with just a laptop. You can see what I am building right now on my projects page.
Frequently asked questions
What free tools can I use to build a python voice assistant?
You can use VS Code for editing, Claude for code suggestions, Gemini’s Whisper model for speech‑to‑text, the macOS say command or any free TTS, and Supabase Edge Functions with the DuckDuckGo Instant Answer API for web search.
How does JARVIS decide when to search the web?
JARVIS checks the chat model’s reply for phrases like "I don’t know" or overly generic answers. If it detects uncertainty, it calls the Supabase function that queries DuckDuckGo and feeds the result back to the chat model for a refined response.
Can I run JARVIS on Linux?
The current version uses the macOS say command, but you can replace it with any free cross‑platform TTS library. The rest of the code – voice recording, Claude integration, Supabase search – works on Linux as long as you have ffmpeg and Python installed.
Do I need to pay for any API to use JARVIS?
All the services I used have free tiers: Claude’s VS Code extension, Gemini Whisper, Supabase Edge Functions, and DuckDuckGo’s Instant Answer API. You can stay within the free limits for a personal project.