← All labs

Local AI Lab

Two small open-source apps that let students chat with an open AI model running on their own computer. Nothing is sent to the internet: prompts, notes and answers stay on the device.

Sign in to download the starter kitLicence: MIT. Each AI model has its own licence; check it before class.

What's inside

Browser app

Streaming chat, four teaching modes, temperature control, “answer from my notes” with cited sources, and a receipt under every answer (model, tokens, speed, sources).

Needs: Python 3.10+, FastAPI, Uvicorn, httpx
Desktop app

Desktop window with streaming chat, modes, temperature, receipt and save transcript.

Needs: Python 3.10+ only
First-steps script

Call the model from Python and measure speed at two temperatures.

Needs: Python only

Labs by week

  1. Week 10 · Lab 7Run the first-steps script with two models; fill in the comparison table (speed, quality on 10 questions, memory, licence).
  2. Week 11 · Lab 8Add one feature to each app (a new mode, dark/light toggle, Bangla labels, copy button, notes upload, token count). Build a 10-question test set for your notes and record correct and correctly cited answers.
  3. Week 12 · Lab 9Add a safe tool (calculator or unit converter) with an allow-list; log every tool call.
  4. Week 13 · Lab 10Run 20 test prompts across users, languages and tricky requests (including hidden “ignore your instructions…” in notes). Record failures and fixes.
  5. Weeks 14–15 · CapstoneTurn the starter into your own app (study helper, club FAQ bot, language partner…) with a test report, a responsible-AI card and a README.

Setup

1. Install Ollama and a model (once per computer)

  • Install Ollama from ollama.com/download (Windows, macOS, Linux). On Linux: curl -fsSL https://ollama.com/install.sh | sh
  • Pull a chat model that fits the computer's memory (see the table).
  • For “answer from my notes”, also pull an embedding model: ollama pull nomic-embed-text
  • Test it: ollama run qwen3:1.7b "Explain a token in one sentence."

2. Run the browser app

  • macOS / Linux: ./run_web.sh — Windows: double-click run_web.bat
  • Open http://127.0.0.1:8000
  • Manual start: cd webapp && pip install -r requirements.txt && python server.py

3. Run the desktop app

  • ./run_desktop.sh (or run_desktop.bat on Windows, or python desktop/local_ai_lab_desktop.py)
  • On Linux, if Tkinter is missing: sudo apt install python3-tk

Which model for which computer

Computer memorySuggested open modelsCommand
8 GB (most school laptops)Qwen3 1.7B or 4B, Gemma 3 1B or 4B, Llama 3.2 3Bollama pull qwen3:1.7b
16 GBgpt-oss 20B, Qwen3 8B or 14B, Gemma 3 12Bollama pull gpt-oss:20b
Lab workstation, large GPU (≈80 GB)gpt-oss 120B or other large open modelsollama pull gpt-oss:120b

Notes for teachers

  • Privacy: with local models, student prompts never leave school devices — helpful for FERPA and state laws in the US and the Personal Data Protection Act in Bangladesh. Do not add cloud services to student builds without checking school policy.
  • Shared server option: one teacher computer with 16–32 GB can serve the class (OLLAMA_HOST=0.0.0.0 ollama serve; students set OLLAMA_URL=http://<teacher-ip>:11434). Use only on a trusted school network — Ollama has no login.
  • Keep the browser app on 127.0.0.1 unless you deliberately share it on a trusted network.
  • AI use: every answer has a receipt and saved chats include an AI-use disclosure line. Follow the course's three AI-use tiers.
  • Model names, sizes and tags change often. Check ollama.com/library and each model's licence before class.

Troubleshooting

“Ollama is not running”Open the Ollama app, or run ollama serve in a terminal.
“No chat models installed”ollama pull qwen3:1.7b (or another model from the table).
“Embedding model … is not installed”ollama pull nomic-embed-text
Answers are very slowUse a smaller model, close other apps, or lower “Longest answer”.
Port 8000 already in usePORT=8010 python server.py