--- title: what conversational ai lacks lab: Quippy, New York studies: conversational personal assistants format: thought dump, lightly edited status: half-published, still arguing claims: none proven. five hypotheses. ---
Most assistants have met you exactly once. It keeps happening.
abstract. Quippy is a small research lab in New York working on conversational personal assistants. These are working notes on what today's conversational AI misses about the people it talks to. Hypotheses, not findings. Some are probably wrong; we'd like to know which.
+------------------------------------------+ | . ... ... ... ..... ...... ... ..... | | ... ..... ...... ....... ....... | | ... ..... ..... .. .. ........ ...... | | ... ..... ....... ... .... ..... .... | | ..... ... .... .. ... .... ... | | | | > can you move my 3pm? | | | | reply: "sure! what time works better?" | +------------------------------------------+ seen: 1 message. 5 words. no yesterday.
+------------------------------------------+ | 3 wks ago "no calls before 10, ever" | | sun kid's fever. better? unclear | | mon 23:10 email to J: drafted, unsent | | tue 08:02 skipped the gym. third time | | 15:00 1:1 with J. the hard one | | | | > can you move my 3pm? | | | | under it: buy me a day? or a way out? | +------------------------------------------+ carried: 3 weeks. none of it typed.
01 memory is not a transcript
Assistants remember in one of two ways: a context window that evaporates when the tab closes, or a list of facts about you that never ages. Neither is how anyone who knows you remembers you.
People remember with decay, weight, and revision. You said you hated running in March. You've run twice a week since June. A friend notices. An index returns March.
fact: "i hate running", stated once, march
observed index friend
mar said it ###### ######
may ran once ###### #####.
jun ran 2x/week ###### ###...
aug ran 2x/week ###### #.....
sep signed up 10k ###### ......
index: "hates running."
friend: "wait, since when?"
H1Useful memory needs time as a first-class field: when something became true, how sure we were, and what has quietly replaced it.
what would change our mind
If people prefer a static profile they can read and edit over one that drifts on its own. That's plausible. A memory you can't inspect is hard to trust, and legibility may beat accuracy.
02 nobody talks in a vacuum
Almost every real request involves someone who isn't in the chat. The same sentence is logistics with one person and a small negotiation with another.
"tell her i'll be late." her is the right move is ---------------------------------------- your manager move it, apologize once your partner call. don't text. your mother say why, or she'll worry J, today don't. ask me first. the prompt is identical in every row.
Assistants are built for a party of one. Everyone else shows up as a string, if at all, and none of them agreed to be in the prompt.
H2An assistant needs a model of the people around you, and a separate model of what it's allowed to know about them. The second is harder and matters more.
what would change our mind
If people simply don't want this. A system that models your relationships can feel like surveillance even when it's right. We treat that as the central design problem, not a footnote. (Yes, this is a footnote.)
03 the turn-taking trap
Chat runs on a strict rhythm: you speak, it answers, then it waits indefinitely. It never cuts in and never goes first. That reads as polite. Mostly it's an interface limit dressed up as manners.
chat, as built
you #### #### ####
it #### #### ####
it never cuts in. it never starts.
someone who's good at this
you ######### ### #####
it ^ (silent) ^
| |
"wait, isn't that "you asked me
your flight?" to nudge you."
People who are good at helping interrupt at the right moment and hold back at the wrong one. The skill is timing, not talking.
H3Generating a suggestion is the easy half of initiative. The hard half is a cost model for being wrong about when.
what would change our mind
If unprompted messages land as noise even when they're correct. Notification fatigue is real and could sink the whole idea. We'd rather learn that early than late.
04 confident about the wrong things
Models sound exactly as sure about your dentist's address as about the boiling point of water. Hedging exists, but it's a verbal tic, not a signal.
source belief then never told ? ? ? ? ? ask told in march ####...... check inferred ~ ~ ~ ~ ~ say so told, current ########## act most assistants: same voice for all four.
The uncertainty that matters day to day isn't about the world. It's about you: what you never said, what you said long ago, what's only a guess from habit.
H4Calibrated uncertainty about the user is under-studied, and probably more useful than calibration on trivia. Saying it briefly, with a source and a date, may be most of the value.
what would change our mind
If visible doubt is just friction. People want answers. The bet is that four words, "from March, maybe stale", cost less than one confident mistake. That bet could lose.
05 the last prompt problem
Whatever you typed last becomes you. Ask one question about divorce law and the next ten answers turn gentle. Ask for short replies once and it stays curt while you draft a wedding toast.
9am precise \ 2pm a parent \ 6pm venting >--> [ last prompt ] 11pm tired / sat someone else / plural in, singular out.
People are plural: sharp at 9am, a parent at 2pm, venting at 6pm, someone else entirely on a Saturday. The model sees one register and overfits to it.
H5An assistant should hold several partial models of a person and be slow to collapse them into one.
what would change our mind
If recency really is the best signal. For the next message, it might be. We suspect it's the best predictor of the next message and one of the worst for the next month.
06 contact
We're a research lab working on conversational personal assistants. Small, based in New York, mostly arguing about the questions above.
If you think we're wrong about any of this, that's the email we want most.
Email ushello@withquippy.com