seenzus’s home screen used to open with a flattering headline: “Handled N things for you today.” The number came out of the operation log, updated daily, and sat in the most expensive spot in the product. We deleted it. Not out of modesty: when we checked the number line by line, it failed in four places. It counted issued commands as “handled,” although the operation log does not know whether a command actually succeeded. It excluded scheduled tasks, even though the section directly below shows exactly those results. Its “today” rolled over at the device’s midnight rather than the home’s own clock. And when the number could not load, the page fell back to “A quiet day so far,” describing a network failure and a genuinely calm home in the same words.
One sentence meant to prove diligence, wrong four ways. That flipped the direction of our worry. Going in, we worried that a proactive Smart Space Agent would miss things. Assembling the home screen taught us the crowded side is the speaking side. Weather wants to mention an umbrella. Learning wants to report a new habit. Device state wants to announce that something went offline. Scheduled tasks want to show last night’s results. Each source, taken alone, has a good reason to speak; put together, they make a notification feed. A notification feed and a wall of device toggles do the same thing to your attention: they hand the “should I care about this” judgment back to you, in bulk.
So across these three months we have spent less effort making seenzus speak than stopping it from speaking. Speaking is a managed act now: one voice at a time; evidence before words; a “no” from you gets corrected first; and when it is not sure, it says so in words instead of dressing the guess up as a percentage. The rest of this post walks through those constraints, then what they cost.
Home carries at most one voice at a time
No source writes to the home screen directly. Weather, learning results, and device state can only submit “I have something to say” to a single arbiter, which decides whether anything deserves your attention right now and, if so, picks exactly one thing. That thing becomes the sole protagonist and the screen enters its busy state. No protagonist, and home is a verifiable line instead: I’ve had a look, nothing needs you. The two states are exclusive, and at most one action waits for you per screen.
The pecking order is fixed: verified urgent risk first, then time-boxed action proposals, then due reminders, and new discoveries last. Within a layer, candidates compete on relevance to the space you are in and on how fresh their evidence is. Losers do not queue. Once you handle the protagonist, this visit goes quiet and nothing backfills; the one exception is an urgent alert that is still valid. We did not build an inbox that pops the next item when you close one, because an inbox turns the home screen into debt. Discoveries also carry a hard quota: one per person per day.
The most direct consequence of this arbitration is that a lot of perfectly harmless remarks get suppressed forever. We accept that.
Speaking has to clear evidence first
A single event is not a reason to speak. The exception is anything you explicitly asked it to watch, a channel you authorized yourself; we come back to that below. A light on once at three in the morning, no coffee brewed today: these enter the record, but one isolated observation is not worth anyone’s attention until it sits alongside longer evidence.
A reminder that made the schedule is re-checked at the moment it would appear. If the window is already shut, there is no close-the-window reminder. If the rain window has passed, the umbrella is off the table. If the upstream forecast has been withdrawn, so has the topic. One rain gets one mention per day; answer “got it” and rain stays off the screen until tomorrow.
A freshly connected space is the harshest case of this rule: seenzus knows almost nothing about it, so it says almost nothing beyond “I’m keeping an eye on things.” For the first two weeks it looks uneventful. We accept this too: enthusiasm without evidence turns into noise within days.
A single no outranks ten yeses
One basic form of speaking is holding up a concrete observation for you to judge, something like “lately the living-room light is often still on after ten,” and you answer right or wrong. Both answers are evidence. They are treated very differently.
“Wrong” takes effect immediately. That observation closes for good, and the same thing will not come back reworded. The learning subject behind it is slotted into the next learning batch for priority review. Before the topic can come up again, seenzus needs materially new behavioral evidence, and even then it waits for the normal cadence. “Right” jumps no queue; it is recorded and digested by an ordinary later batch. The asymmetry is deliberate: something said wrong has to be fixed promptly, and being confirmed does not justify displacing unsolved problems.
The other rule matters just as much: silence is not a no. A card you walk past, ignore, or let expire is never converted into “he dislikes these reminders.” No response means no response. The product does not invent a position for your silence.
The internal confidence score never becomes a number on screen
seenzus does keep confidence internally: every learned rule carries a belief score that rises and falls with feedback, and learning uses it to rank and prune. That number appears on no user-facing surface. No “80% sure,” no “12 out of 14 nights,” no “observed for 14 days.” When it is not sure, it says “waiting for you to confirm,” or “let me watch a little longer,” or it stays silent.
The reason is that a percentage, here, is a performance. “86% hit rate” reads like honesty, but it hands the judgment right back to you: whether that number is high enough to trust is now your math to do. We want it to own “I’m not sure” in words. The numbers you do see describe the world: sixteen devices, seven of them on, idle for two hours. Those are facts it observed, not grades it gave its own guesses.
This line has a cost of its own. Some users will want to see exactly how sure it is, and will read the absence as opacity. Our current answer: every sentence it says can be checked against the world, and that is closer to transparency than putting a price tag on a guess.
The top priority layer has exactly one source today
The first layer of the arbitration is verified urgent risk. Today that layer has a single tenant: the devices you marked as important, by hand, on the device page. A flagged device has to stay unreachable for five continuous minutes before it qualifies as an alert and takes the protagonist slot, and the copy claims only what was verified: contact lost, cannot confirm it is running. It does not guess a power cut and does not guess trouble. Beyond that, the layer is empty. No risk seenzus infers on its own has earned a place there, and ordinary bad weather is never promoted into it to fill space. The only things allowed to interrupt you at top priority are things you pointed at yourself.
This is where quiet presents its bill. A system built to avoid false alarms leaves the risk of missed ones on the table. People tolerate the two errors asymmetrically: one unsaid thing is usually forgiven, while a few unnecessary interruptions in a row can finish off trust. That is why we guard harder against false alarms. But that is a decision about odds, and the question of responsibility stays open: when a user says “interrupt me less,” what exactly are they giving up? If something that should have been said was not, whose failure is that? We do not yet have an answer we are satisfied with.
Back to the deleted headline. Its problem can be stated more precisely now: it used a number to prove the product’s presence, which is exactly the habit we are asking seenzus to break. What these three months settled is this: in this product, speaking is a scarce act that has to be applied for with evidence. When the home screen is quiet, most of the time that is the constraint working. Whether this quiet will one day fail at a moment that mattered, only longer life in real spaces can tell. We are leaving that question on the table.
seenzus 的首页有过一句很讨喜的大字:「今天替你办妥了 N 件事」。数字从操作记录里数出来,每天更新,放在打开产品看到的第一个位置。后来我们把它删了。不是出于谦虚,是因为逐项检查时,这个数字在四个地方站不住:它把「发出了指令」算成「办妥了」,而操作记录并不知道每条指令最后是否成功;它不统计定时任务,可首页下方那一栏「替你办好的」摆的恰恰是定时任务的成果;它的「今天」按设备所在时区的午夜切分,和这个家自己的钟对不上;取不到数字时,页面垫一句「今天很安静」,把加载失败和真的无事说成了同一件事。
一句用来证明勤快的话,四处失实。这件事把我们对「主动」的担心掉了个头。做一个会主动开口的 Smart Space Agent,我们起初操心的是它看漏什么;等到把首页真正拼起来才发现,拥挤的是想说话的那一侧。天气想提醒带伞,学习想汇报新发现的习惯,设备状态想报告有东西掉线,定时任务想展示昨晚的成果。每个来源单独看,开口都有理由;放在一起,首页就成了通知流。通知流和一屏摆满开关的控制面板,对用户注意力做的是同一件事:把「要不要管」的判断成批地留给人。
所以这三个月,我们花在「让它开口」上的功夫,少于花在「拦住它开口」上的。现在的 seenzus 里,开口是一个被管理的动作:同一时刻只有一个声音;说之前要过证据;被否定一次,纠正优先;说不准的时候用话承认,不拿数字装可信。下面把这几条约束逐条摊开,最后说它们的代价。
首页同一时刻只允许一个声音
想开口的来源,谁都不能直接往首页写东西。天气、学习结果、设备状态,都只能把「我有一件事想说」提交给同一个仲裁点,由它判断此刻有没有值得用户看的事;有,就只选一件。选中的事成为首页唯一的主角,页面进入「有事」状态;没有主角,首页就是一句可以核对的交代:我替你看过了,目前不用操心。两种状态互斥,一屏至多一个等你处理的动作。
主角的挑选顺序是固定的:已证实的紧急风险最先,其次是有时效的操作提议,然后是到点的提醒,最后才是新发现的建议。同一层里的候选,再按与你当前所在空间的相关性和证据的新鲜程度比较。落选的候选不排队:你处理完主角,这次打开就归于安静,后面的不会补上来;只有仍然有效的紧急警报例外。我们不做「关一条冒一条」的信箱,信箱会把首页变成还债。新发现类的建议另有一条硬配额:同一个人,同一天,至多一条。
这套仲裁最直接的后果,是很多「说了也无妨」的话被永久地憋住了。我们接受这个结果。
开口要先过证据这一关
单个事件构不成开口的理由;你明确点名让它盯着的东西是例外,那条通道是你自己授权的,后面会讲到。深夜某盏灯亮了一次,今天咖啡机没有开,这些都会进入记录,但一条孤立的观察不值得惊动任何人;它要和更长时间里的其他证据放在一起,才谈得上意义。
已经排上的提醒,在真正露面前的那一刻还要重新核对一遍。窗户已经关上,就不存在「关窗」的提醒;降雨的时间窗过了,伞的提醒跟着作废;要是上游把预报整个撤了,雨的话题就撤。同一场雨一天只说一次,你回一句「知道了」,当天它不再提。
刚接入的空间是这条规则最严苛的情形:seenzus 对它几乎一无所知,于是几乎什么都不说,只留一句「我先替你看着」。它头两周显得没什么动静。我们接受这一点:没有依据的热闹,过不了几天就会变成打扰。
否定比确认更需要立刻处理
seenzus 开口的一种基本形态,是把一条具体观察拿给你判断,比如「最近一段时间,晚上十点后客厅的灯经常还亮着」,你只回答对或不对。两种回答都是证据,待遇却完全不同。
「不对」立刻生效。这条观察永久关闭,同一件事不会换个说法再来;它背后的学习课题会插进下一批学习,优先重新审视。要再谈同一个话题,得先有新的行为证据,并且仍要等正常的开口节奏。「对」不插队,它被记下,由之后的批次正常消化。这个不对称是有意的:说错的话必须尽快改,被确认的判断不值得挤掉还没解决的问题。
另一条规则同样要紧:沉默不算否定。一张卡片被路过、被无视、到期消失,都不会折算成「他不喜欢这类提醒」。不回应就只是不回应,产品不替沉默编造立场。
内部的把握度永远不变成界面上的数字
seenzus 内部有把握度:每条学到的规则都带一个随反馈升降的信念值,学习拿它来排序和取舍。但这个数字不出现在任何用户界面上。不写八成把握,不写「12/14 次」,不写「已观察 14 天」。说不准的时候,它说「待你确认」,说「再看看」,或者选择沉默。
理由是,百分比在这里是一种表演。「命中率 86%」读起来像坦白,实际把判断责任又推回给用户:数字够不够高,值不值得信,你自己算。我们要它连「说不准」也用话承担起来。界面上出现的数字只描述世界:十六台设备,七台开着,闲置两小时。那些是它看到的事实,不是它给自己的猜测打的分。
这条线也有代价:会有用户想看看它到底多有把握,并因此觉得它不够透明。我们目前的答复是,它说出的每句话都可以核对,这比一个猜测的评分更接近透明。
风险那一层今天只有一个来源
前面说仲裁的第一层是「已证实的紧急风险」。这一层今天只有一个来源:你在设备页亲手标记的重要设备。被标记的设备持续失联满五分钟,才有资格以警报的身份成为首页主角,文案也只说核实过的部分:失联了,无法确认还在工作。它不猜断电,不猜出了什么事。除此之外,这一层是空的:seenzus 自己推断出来的任何「风险」,今天都还没有资格住进来;普通的坏天气,也永远不允许被拔进这一层充数。有资格用最高优先级打扰你的,只有你亲自点过名的东西。
安静的代价就在这里。一套处处防误打扰的机制,天然把漏报的风险留在桌面上。人对这两种错误的容忍并不对称:一次该说没说,多半会被原谅;连续几次不该说而说,信任很快见底。我们据此把误报防得更严。但这只是在概率上占了便宜,责任的问题还在:用户说「少打扰我」的时候,他愿意让渡的究竟是什么?真错过了一条本该说的话,后果算谁的?这道题我们还没有让自己满意的解法。
回到开头那句被删掉的大字。它的问题现在可以说得更准:它想用一个数字证明产品的存在感,而这正是我们要求 seenzus 戒掉的东西。三个月里定下来的是这一条:在这个产品里,开口是要凭证据申请的稀缺动作。首页安静的时候,多数是这套约束在正常工作。至于这份安静会不会在某个要紧的时刻失职,得靠真实空间里更长时间的使用来回答。这个问题我们留在桌面上。