
The Chatterbox card
Make a believable chatbot without an LLM (like ChatGPT or Claude)! Markov chains, keywords, etc… you can use any training data!
Make a chatbot that’s believable enough to hold a conversation, without using an AI model like ChatGPT or Claude!
What counts
The chatbot has to come from your own code. Rule-based replies, keyword matching, Markov chains, digging through a pile of old conversations for a good reply… all of those count, and you can train them on any data you like.
The one thing that doesn’t count is an LLM. Calling ChatGPT (or any other API like it) is out, and so is running one on your own machine. A really tiny neural network you train yourself is totally fine though! If you’re not sure about something, ask in #crescent-help before you build on top of it.
Meet ELIZA
The first chatbot people really got attached to was ELIZA, written by Joseph Weizenbaum at MIT in the mid-1960s. Its most famous script pretends to be a therapist, and nearly everything it does is spot a keyword, flip “I” and “you” around, and ask a question back.
You: I’m worried about my exams
ELIZA: Why are you worried about your exams?
People still told it their secrets! Weizenbaum was honestly a bit alarmed by how quickly everyone decided it understood them. There’s plenty of versions online you can try.

A conversation with one of the many ELIZAs out there. Screenshot via Wikimedia Commons, public domain.
Building your own ELIZA is the easiest thing that counts. You need a list of patterns, a few replies for each, and something to say when nothing matches. That’s an evening of work, and then you get to spend the rest of the week making it less obviously a robot.
Making it believable
This is the fun part! A bot that answers everything with “Interesting, tell me more.” gets old after about four messages. Some things that help a lot:
- Remember stuff. If someone tells it their name, or their dog’s name, bring it up later. One remembered fact makes a bot feel way smarter than it is.
- Give it a personality. A grumpy pirate, a nervous wizard, your cat… A character gives you an excuse for weird answers, and it’s much more fun to talk to than a polite assistant.
- Dodge gracefully. Your bot won’t understand most things, and that’s fine! Changing the subject or asking something back beats a confused non-answer. ELIZA was built almost entirely on this trick.
- Fake the typing. A short pause before replying, a bit longer for longer replies, makes a surprising difference.
Going further
- Markov chains. Feed in a big pile of text, count which word tends to follow which, then walk through those counts to make new sentences. They ramble, but they ramble in the voice of whatever you fed them, which can be really funny. Try one trained on a book from Project Gutenberg!
- Retrieval. Take a pile of real conversations (the Cornell Movie-Dialogs Corpus has around 220,000 exchanges from movie scripts), find the line closest to what the user typed, and reply with whatever came after it. Look up TF-IDF for a simple way to measure “closest”.
- Intents. Sort messages into things the user might want (“tell me a joke”, “what’s the weather”, “who are you?”), pull out the useful bits, and keep track of where you are in the conversation. Lots of support bots worked like this before LLMs came along.
- A tiny neural network. Train a little one yourself, from scratch, on your own data. It won’t be ChatGPT, and that’s the point! Getting something that small to say anything sensible at all is a real achievement.
- Mix them! Rules for the things you want to get exactly right, and a Markov chain or retrieval for everything else.
About training data
Any data is fine, but if you use chat logs, stick to your own messages or ask your friends first. Nobody wants their DMs showing up in someone else’s bot!
Shipping one
- Demo: a web page where anyone can chat with it is ideal. If it’s a Slack or Discord bot, make sure it’s actually running and tell us where to find it.
- Screenshot: a conversation where it’s on its best behaviour. We know you picked a good one, that’s allowed!
- README: how it works and where the training data came from. A couple of real transcripts are great to include, especially one where it gets completely confused.
