# Bhekani.com — Full Blog Content > All published blog posts from bhekani.com, formatted as markdown for LLM consumption. ## We're building the culture of AI work right now - **Date**: 2026-03-16 - **Summary**: Right now, every way we talk about AI at work is helping define what ownership, competence, and judgment look like in the AI era. We should be more deliberate about that. - **Tags**: ai, musings, culture, work - **URL**: https://bhekani.com/posts/were-building-the-culture-of-ai-work-right-now/ Right now, across the industry, engineers are figuring out how to work with AI. How to use it for brainstorming, for drafting, for research, for code. How much to trust it. When to override it. How to fold it into workflows that already existed before any of this showed up. That part is obvious and everyone is talking about it. What fewer people are paying attention to is the other thing that's happening at the same time. We're also building a culture around how we collectively perceive, talk about, and judge work that involves AI. Every casual phrase we attach to the things we share is helping set norms. Norms about ownership. About what counts as real work. About whether AI involvement is something to apologize for or something to own. Most of this culture-building is happening by accident. That's exactly why it matters. ## Everyone is navigating this differently If you've shared AI-assisted work recently, you've probably thought about how to frame it. And there's no established playbook. The norms are genuinely unsettled. Some people say "Claude helped with this" because their workplace expects disclosure. Some say it because they want to be transparent about their process. Some say it because they're excited about what AI can do and want to show it off. Some say it because the norms still feel unsettled, and hedging feels safer than ownership while the social rules are still being written. All of that is understandable. Nobody taught us how to talk about this. But some of those improvisations are becoming habits, and some of those habits are shaping the culture in ways worth examining. ## The pattern I keep noticing I keep seeing experienced engineers share work with a little disclaimer attached to it. "This was generated by Codex." "Claude helped with this." "Most of this came from AI." Sometimes the phrasing is transparent and informative. But often, something else is going on. The disclaimer isn't really explaining the process. It's creating distance between the person and the work. Quietly asking the reader to lower their standards. It can also run in the opposite direction. I've seen senior engineers mention AI not to hedge, but almost to flex. The subtext is less "don't judge me too hard" and more "I'm so secure in my own competence that I can openly admit AI helped." That's not defensive. But the sentence is still doing social work around the author rather than clarifying the artifact. Whether the motive is insecurity, transparency, status, or genuine uncertainty, the phrase "AI helped with this" often ends up doing cultural work beyond the literal words. Is AI use something to apologize for? Is it a way to lower standards? Is it an excuse to stop thinking? Is it a workflow detail? Is it evidence of modern competence? The answer depends on the norms we set now, and those norms are being built out of a thousand little repeated moves that nobody bothers to question. ## We already knew how this worked The easiest way to see what I mean is to remove AI from the story. In the pre-AI world, I might work on something with a coworker. Maybe I bounced ideas off a peer. Maybe a junior engineer drafted the first version. Maybe I pair programmed with Mike and he was typing for most of the session. When I go to post the result, it sometimes makes sense to mention the other person. Credit them for a real contribution. Acknowledge collaboration. That's normal. But what I would not do is mention them as a way of hedging responsibility. I would not say "Mike wrote most of this code" in a tone that really means "so if this is bad, direct some of that at Mike." I would not say "Sarah came up with a lot of this" in a tone that means "please don't judge me too hard." That would be weak. I'm still the one posting it. I'm still the one saying: this represents something I'm willing to put forward. Engineers have always worked collaboratively. We use docs. We use Google. We use Stack Overflow. We use examples. We ask coworkers. We brainstorm in chats. We sketch a bad version and get feedback. We review each other's pull requests. We pair. We copy a pattern from another codebase. We have one person navigate while another types. We argue through trade-offs. We reformulate an idea six times before landing on the version that works. None of this has ever threatened the basic idea of ownership. The artifact belongs to the people who stand behind it. AI is now part of that same family of collaboration. Sometimes it plays a small role, sometimes a large one. Either way, I think its involvement should be treated the same way we treat other forms of collaboration: mention it when it's relevant, but don't use it to smuggle in a disclaimer on your own responsibility. ## Human credit and AI "credit" are different things The coworker analogy works, but it also reveals something deeper. When I mention a human collaborator, there's often a real moral and social reason for it. They deserve recognition. They may care whether their work gets erased. There's a relationship there, professional visibility, fairness. If I pair programmed with Mike and his contribution was substantial, there's a genuine social reason to say so. Mike is a person. He has feelings. He's building a career. Not crediting him would be a kind of erasure. When I mention Claude or Codex, that logic doesn't carry over. Claude doesn't care. Claude is not hoping this work helps it get promoted, and it's not a colleague whose effort I'm morally obliged to recognize. So when AI "credit" looks like human credit but doesn't serve the same purpose, it's worth asking what purpose it is serving. Often, it signals "this was not fully me," which easily turns into "please calibrate your judgment accordingly." ## Where I think we should go: say what you mean Mentioning AI is sometimes exactly the right thing to do. The goal isn't silence about AI use. Silence has its own problems, and a culture where everyone quietly uses AI but never talks about it would be a different kind of bad. The goal is precision. Maybe you want to show that these tools are no longer toys. Say that. "I used Claude heavily here because I wanted to show what this kind of collaboration can produce." That's a real statement. Or maybe you're explaining workflow: "This was vibe-coded." "The model generated most of the first pass and I refined it." Those are useful descriptions. Or maybe you're communicating the status of the artifact: "This is a quick PoC, not production code" is clearer and more honest than "AI wrote this lol." But be precise. "I'm highlighting AI capability" is not the same as "this is not production-quality." "I collaborated heavily with Claude" is not the same as "please lower the standard you apply to me." Those are different claims. Say the one you actually mean. A lot of people say "AI wrote this" when what they actually mean is something else. This is rough. This is a PoC. I'm showing what AI can do. I moved quickly and didn't polish every edge. The model made a big contribution. I'm excited that AI is capable of participating meaningfully in this kind of work. All of those are valid things to say. So say them. The moment "AI wrote this" becomes a vague proxy for all of them, it stops being informative and starts being social insulation. That's when the phrase starts doing cultural work beyond what anyone meant by it. And yes, sometimes AI provenance is itself useful information. AI has specific failure modes. Reviewers might want to know. Fair enough. But if what you're really trying to say is "I'm not sure this is fully reliable," the answer is more review, not a disclaimer. The disclaimer doesn't fix the problem. It just transfers the burden to the reader. ## The norms are still wet cement We are still early in AI-era work. The norms are still forming. They will harden. And they're being shaped right now by the small, repeated ways we talk about AI involvement in our work. If vague AI attribution becomes the default, I think we risk teaching people two bad lessons at once. The first: that AI-assisted work is something you should subtly distance yourself from. AI is already too useful and too normal for that line to hold. The second: that the machine's involvement dissolves your responsibility. That one teaches people to stop thinking, stop editing, stop reviewing, and stop owning. This matters because habits of speech become habits of judgment. If we normalize vague AI disclaimers, we teach people that AI use weakens legitimacy, that responsibility becomes blurry once a model is involved, and that precision about process is optional. Those are bad norms, and they will compound. There's also something that happens to the person doing the work. If every time you use AI you mentally frame the output as something slightly outside yourself, you make it easier to skip the hard part, which is judgment. You become a courier instead of an editor. A presenter of outputs instead of an owner of decisions. The real value of AI is not that it replaces your responsibility. It changes where your effort goes. Less raw drafting, more evaluation, direction, and judgment. If people don't internalize that, they will use AI in the laziest way possible and wonder why the results feel hollow. And we are in the middle of redefining what competence looks like. In the old world, people could pretend the highest form of competence was doing everything manually. That was never fully true, but AI makes the myth much harder to defend. The question now is: what counts as real skill? I think the answer is sound judgment over increasingly powerful tools. Directing, evaluating, integrating, and deciding. The competent person is the one who can guide AI toward something worth standing behind, not the one who avoids it. The healthier norm is simpler. AI is a legitimate collaborator, its contribution can be large, disclosure should be precise, and judgment stays with the human. ## Own the work If you think something is too undercooked to be associated with you, don't post it. If you think it's worth posting, own the level at which it should be judged. Ownership doesn't mean pretending everything is finished. It means standing behind what you chose to share. If you read the work, approved it, and chose to post it under your name, own it. Whether the AI contributed five percent or ninety-five percent, what matters is that you adopted it. Treat AI more like a coworker and less like a contamination warning. --- ## The real bottleneck in your agentic workflow is you - **Date**: 2026-03-14 - **Summary**: When an AI agent stops to ask you something, it's usually not a reasoning failure. It's a context-access failure. The fix isn't a smarter model. It's better systems for capturing the judgment trapped in your head. - **Tags**: ai, technical, agents, memory - **URL**: https://bhekani.com/posts/the-real-bottleneck-in-your-agentic-workflow-is-you/ I spend a lot of time coding with AI agents, and a pattern keeps showing up that I think most people misread. The agent gets stuck. It stops and asks me something. I answer. It carries on perfectly well. The standard interpretation is that the model wasn't smart enough. But look at what actually happened: it asked one question, got one missing piece, and continued. If the model were incapable of doing the work, my answer wouldn't have unblocked it. I would have needed to reason through the whole thing myself. That's not what happens. The model hits a boundary where some piece of context lives in one place only: my head. The interruption is not a reasoning failure. It's a context-access failure. Once you see it that way, a different design problem appears. The question stops being "how do we make the model smarter?" and becomes "how do we reduce the number of times the system has to query the human for context that could have been available already, or captured once and reused later?" When an agent interrupts you mid-task, it is doing something simple. It is querying you. Almost like an API call to a system it can't fully inspect. The model doesn't have access to the assumptions in your head, your project-specific judgment, your preferences, your mental map of the codebase, the decisions you made three weeks ago that still shape how the work should be done today. So it asks. I think human-in-the-loop agent systems are best understood as a cache architecture. Your brain is the origin server. The agent's memory, documentation, skills, prior examples, and operational rules are the cache in front of it. The goal is to avoid hitting that origin for the same class of question over and over again. ## What's actually stuck in your head Most people assume that what's stuck in the human's head is project facts. That matters, but it's too narrow. When you step in to answer an agent's question, you're not always giving it information. Often, you're giving it judgment. Sometimes it's concrete facts: this service behaves differently in staging, this endpoint is technically deprecated but another team still depends on it, this customer workflow matters more than the internal abstraction. Sometimes it's preferences about how you or your team works: prefer explicit if statements over ternaries, keep orchestration logic out of the route handler, don't introduce another config format unless absolutely necessary. Sometimes it's principles: fix root causes rather than patching symptoms, prefer reversible changes when ambiguity is high, don't optimize away clarity in core flows. Often it's heuristics, compressed judgment from experience: if something touches billing, auth, or migrations, slow down. If there are two plausible interpretations, choose the one with the safer rollback story. If a failing test looks flaky, check shared mutable state before touching production code. It could be decision rules you apply repeatedly: if the change touches a shared interface, prefer the adapter layer before changing domain types. If the request is underspecified and low-risk, act with the most unsurprising default. If the action is hard to reverse, stop and ask. It could be workflow steps about how work actually gets done: before changing this pipeline, inspect the event shape at all three boundaries. Before merging a migration, assess lock risk and reversibility separately. Or it could be exceptions, things that look fine generically but are wrong here: don't reuse that helper, it has side effects nobody expects. Don't copy that older pattern, it exists for historical reasons. This is what makes the problem interesting. The missing thing is usually not an answer. It's a decision policy. And decision policies can often be externalized enough to reduce future interruptions. ## Why "just document everything" doesn't work The obvious response is: fine, so write it down. Document your facts, your preferences, your heuristics, your decision rules. Problem solved. Two things make that harder than it sounds. First, you can't write down everything. There's always tacit knowledge. Knowledge you don't realize you have. Knowledge you have but can't articulate yet. Pattern recognition from years of experience that doesn't compress into a rule. I'm not claiming that every part of human reasoning can be exported into a machine-readable system. But the opposite mistake is more common and more damaging: assuming that because not everything can be captured, there's no point capturing the recurring parts. When people complain that "the model keeps asking me things," they usually don't mean the task is irreducibly human. They mean there's a recurring class of missing context that nobody has turned into shared memory, rules, or a skill. That's a solvable problem. Second, you usually don't know in advance which details will matter. You don't know which ambiguities will recur. Nobody sits down one day and transcribes their working knowledge into a neat operating manual. And if they tried, most of it would be low-signal junk. I've [written before](/posts/your-agents-md-is-probably-hurting-your-agent/) about how dumping everything into a static context file often makes agents worse, not better. The instruction budget is real — every irrelevant rule competes with the agent's own reasoning. Static documentation is necessary but not sufficient. The better pattern is dynamic capture. Let the interruptions themselves show you where the cache misses are. Every time the agent stops and asks a meaningful question, you've learned something. You've found a place where live human context was still on the critical path. That interruption is diagnostic. It tells you: this class of work still depends on private knowledge. This policy was never externalized. This ambiguity was predictable. This decision pattern should be reusable. Instead of trying to precompute everything, you let repeated friction reveal what matters. A lot of useful knowledge only surfaces at decision time. You don't know what should be a rule until you notice the agent getting stuck on it. You don't know what your own private policy actually is until you hear yourself explain it. ## Store the policy, not the answer This is the part that matters most. If the agent asks "should I use option A or option B?" and you say "use option B," that unblocks the task but doesn't improve the system. The answer is too local. What matters is the policy behind the answer. Why option B? Because option A would create hidden coupling? Because option B is more reversible? Because this area is shared infrastructure and your team optimizes for clarity over abstraction in this layer? The reasoning is the reusable part. Compare these: A bad memory entry says: "use option B here." A better one says: "when a change touches shared interfaces and one option increases hidden coupling, prefer the option that preserves local isolation, even if it is less elegant." A bad memory entry says: "don't use that helper." A better one says: "avoid reusing helpers with hidden side effects in request-critical paths. Prefer explicit local logic if the helper obscures control flow." A bad memory entry says: "ask before doing schema changes." A better one says: "for schema changes, stop and ask when rollback is unclear, lock risk is unknown, or downstream consumers are not visible from the current repo." The second version generalizes. That's what a cacheable human judgment artifact looks like. The parallel to mentorship is hard to miss. When a junior engineer asks the same kind of question repeatedly, the right move isn't to answer the immediate question each time. It's to surface the underlying principle. "In situations like this, prefer X over Y because of Z." Over time, they internalize the policy. The number of interrupts drops. The quality of independent judgment goes up. A good agent system should compound in exactly the same way. Each interruption should make the next one less likely. If it doesn't, you're paying the same cognitive tax every time. The system should also be rewriting your natural, conversational explanations into something operational. Not vague advice like "be careful with migrations" or "use good judgment" or "think about edge cases." Usable guidance like: "Before any migration, inspect current schema shape, estimate lock risk, ensure reversibility, and avoid combining schema and data backfill changes in one deploy unless explicitly justified." That's how hidden judgment becomes executable. ## The cache architecture Once you see all this, the cache metaphor becomes surprisingly complete. You don't repeatedly query the most expensive, highest-latency component in a system if you can front it with a well-designed cache. A good agent architecture should check its own task context first, then explicit rules and project docs, then prior decisions and examples, then extracted heuristics and skills. Only after all of that should it call the human. The sub-concepts map cleanly. A cache hit is when the agent finds a relevant rule or prior decision and proceeds without asking. A cache miss is when the needed context isn't available, so the agent interrupts the human. Cache warming is when you proactively add rules and examples for things you know will come up. Read-through caching is when the agent asks you once, then stores the useful part so future cases are served from memory. Semantic caching is when the next case isn't identical but similar enough that the same policy applies. The hard parts of caching map too. Invalidation matters because old rules go stale. Teams change, codebases change, preferences shift. Guidance that was right six months ago may be wrong now. This connects to something I explored in [building memory systems that forget](/posts/cognitive-memory-for-ai-agents/) — memory that treats everything as equally permanent is memory that eventually drowns in noise. Scope matters because some knowledge is global, some repo-specific, some feature-specific, some only valid for the current task. And some guidance should degrade in confidence over time. A bad cache fails in two opposite directions. It either forgets too much and keeps hitting the human for things it should already know, or it remembers too aggressively and applies outdated rules where they no longer fit. The hard problem isn't "store more stuff." It's "store the right thing, at the right level of abstraction, with the right scope and retrieval behavior." There's a related question the system should be asking: did I actually need to interrupt at all? Before asking, it should check whether it already has a relevant rule, whether it has similar prior examples, whether there's a safe reversible default, whether it's facing true ambiguity or just uncertainty. Some agents over-ask not because context is missing but because they're timid. You want learning, but you also want a bias toward acting when the downside is low and the move is reversible. ## What this reframes All of this points to an uncomfortable conclusion. A lot of what we casually call "AI weakness" is really operator knowledge that has never been externalized. The human carries around facts, heuristics, interpretations of what "good" means in this particular project, memories of prior failures. None of that is available to the system until the human speaks. Reasoning quality still matters. Not all interruptions are avoidable. But a surprising amount of human oversight is the injection of missing context and missing decision policy. Once you see that, you stop thinking only about bigger models, better prompts, or longer context windows. You start asking: what classes of interruption keep recurring? What hidden context caused them? Can it be represented in a reusable form? Can we reserve live human queries for the cases that genuinely need them? This idea has roots in older work on tacit knowledge and externalization, the notion that valuable know-how lives inside people until it gets turned into shareable artifacts. There's also a growing body of work on agent memory and cognitive architectures that tries to classify different kinds of stored knowledge. What I'm adding here is a systems-level framing: an agent interrupt is often a cache miss against human-held context, and the right engineering response is to treat each interrupt as a chance to warm the cache. A simple rule falls out of that: every meaningful interruption should produce either a one-time answer or a reusable artifact. If the answer is too local, it stays local. But if the same class of question is likely to come back, the system should capture the policy behind the answer and make it available next time. That's how the workflow compounds. That's how you reduce interruption frequency without pretending the human is unnecessary. Human attention is expensive, slow, breaks flow, and doesn't parallelize. Use it where it matters: novel decisions, genuine ambiguity, high-risk actions, policy changes, edge cases outside the known patterns. Build everything else so that the system serves itself from the cache before it reaches for the human. The best agent workflows don't try to remove the human. They put a smart cache in front of the human. --- ## What does it actually mean to run AI in production? - **Date**: 2026-03-10 - **Summary**: Your LLM can fail silently. No crash, no error - just worse answers reaching users while your logs show 200 OK. That's one of several ways LLMOps breaks from traditional MLOps. Here's what actually changes when your model is an API call. - **Tags**: ai, technical, mlops, llmops - **URL**: https://bhekani.com/posts/mlops-vs-llmops/ Every software engineer can call an API. Wire it up to an inference endpoint. Ship it. So if every engineer can do this... does that make every engineer an AI engineer? What exactly _is_ an AI engineer? I've been digging into this question, partly to organize my own thinking. There are two ways people tend to get it wrong. The first: AI engineering is just calling an API. Wire it up, ship it, done. The second, more sophisticated version: AI engineering is machine learning, so the operational discipline is MLOps - the mature practice of training, deploying, and monitoring ML models. Both framings miss something. When you're building on top of foundation models you don't own, the operational problems are different enough that they've earned their own name: LLMOps. And understanding why requires understanding what breaks after deployment. Chip Huyen, who literally wrote the book on this (her O'Reilly title [AI Engineering](https://www.oreilly.com/library/view/ai-engineering/9781098166298/) is the closest thing the field has to a reference text), puts it plainly: _"It's easy to make something cool with LLMs, but very hard to make something production-ready with them."_ Gartner [found](https://www.gartner.com/en/newsroom/press-releases/2024-10-22-gartner-says-more-than-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025) that more than 30% of generative AI projects were abandoned after proof of concept by end of 2025. The demos worked. The operations didn't. ## Traditional MLOps: a quick refresher In traditional machine learning, you train a model on labeled data. You collect a dataset, label it, train a classifier or regressor, evaluate it, deploy it, and monitor it for drift. For a fixed model and environment, inference is typically deterministic - same input, same output. You measure success with hard metrics: accuracy, precision, recall, F1 score. If performance degrades, it's usually because the data distribution shifted. Your users started behaving differently than your training data expected. You retrain on fresh data and redeploy. The key thing: **you own the weights.** The model is a file you trained. You control the full change surface - the model, the features, the serving infrastructure. If something changes, it's because someone on your team changed it. The pipeline is well-understood: data collection → labeling → training → evaluation → deployment → monitoring. MLOps tooling exists to automate each step. It's more mature and better understood than what we have for LLM systems. The assumption running through all of it: you own the thing that determines behavior. That assumption doesn't survive contact with LLMs. ## The shift: you're orchestrating, not training You're not training the base model. You're calling an API. This is specifically about that context - teams building on hosted foundation models rather than training or self-hosting their own. The unit of work shifts from "model" to "system": your product is now a combination of prompts, retrieval pipelines, tool integrations, guardrails, and routing logic. Each with its own lifecycle and failure modes. OpenAI, Anthropic, or Google can change model behavior under you with no notice. Your prompts that worked perfectly on GPT-5 might break when the provider updates to GPT-5.2. No code changed on your side. If MLOps is building an engine, LLMOps is wiring together an engine you bought - with custom fuel lines, intake filters, and exhaust monitoring. You didn't build the engine and you can't open the hood, but you're responsible for making the car drive reliably. ZenML's analysis of [1,200 production deployments](https://www.zenml.io/blog/what-1200-production-deployments-reveal-about-llmops-in-2025) found that software engineering fundamentals, not frontier models, remain the primary predictor of success. The engine matters less than how well you've wired it together. ## What breaks after deployment The MLOps instincts that serve you well - deterministic outputs, owned infrastructure, training-time costs, test suites with right answers - break in different ways depending on where you look. Here's where each one fails. ### Silent quality degradation The scariest failure mode in LLM systems is the one that doesn't look like a failure. No crash. No error. Just worse answers. Unlike traditional ML where data drift triggers measurable metric changes, LLMs make silent quality failures especially common - correctness is fuzzy, users often can't verify answers, and there's no ground truth to alert against. Hallucinations reach users without triggering anything in your monitoring. The model confidently makes something up, the user doesn't know enough to question it, and your logs show a successful 200 response. According to LangChain's [State of AI Agents](https://www.langchain.com/state-of-agent-engineering) report (1,340 respondents), **32% cite quality as their number one production barrier.** Not latency, not cost — quality. ### Non-determinism Here's something that still catches people off guard: the Thinking Machines team ran [1,000 identical completions at temperature=0](https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/) and got **80 unique outputs.** Same prompt. Same parameters. 80 different answers. Traditional ML systems can introduce their own variability - recommenders, stochastic pipelines, GPU-level numerical differences. But LLM variability is different in degree and in kind: it's user-visible, product-relevant, and present even when you've explicitly asked for deterministic output. That's what makes it a distinct engineering problem. It breaks the foundational assumption behind most of software engineering's quality tools. Unit tests work because the same input produces the same output - you write a test, it passes or fails, done. That contract doesn't hold with LLMs. You can't write a test that says "given this question, return this answer" because the answer will vary. The whole testing paradigm shifts from pass/fail to probabilistic - which is why the field landed on evals rather than tests. You're not checking whether the output is correct, you're measuring whether it's good enough, often by using another LLM as a judge. That's a different engineering discipline, not just a different tool. Instead of accuracy and recall, you track helpfulness, coherence, latency, hallucination rate, and cost per query. The same shift applies to monitoring: you're not alerting on errors, you're measuring quality distributions over time and watching for drift. Evaluation itself becomes an LLM task - you use one model to judge another model's outputs. It's one of the most practical approaches at scale, especially when combined with task-specific checks and human calibration. I wrote about a related angle in [testing specific AI behaviors](/posts/your-ai-is-confidently-wrong), where even measuring something as simple as "does the model push back on nonsense" requires elaborate benchmarking. ### Cost spiraling In traditional ML, compute costs are frontloaded during training. You pay once, then inference is cheap. Cost is a problem you solve before deployment, not one that compounds with every user request. With LLMs, cost is a continuous operational variable. Every API call costs money, and on most providers output tokens cost significantly more than input tokens - the exact multiplier varies, but it's enough that a poorly optimized prompt can meaningfully inflate your bill on every single request, indefinitely. The mental model of "we paid for the infrastructure, now it runs" doesn't apply. The strategies that work: model tiering (cheap models handle the bulk of routine tasks, expensive models handle the edge cases that actually need them), prompt caching (most providers offer up to 90% reduction on cached input tokens), and [smart routing between models](/posts/building-an-llm-model-router-lessons-from-the-wild). Applied together, these compound to significant savings. None of this exists in traditional MLOps. There's no "prompt cost" in a random forest. ### Latency and reliability Traditional software latency is measured in milliseconds. You cache aggressively, optimize queries, and if something takes more than a second you investigate. That intuition breaks with LLMs. API calls take seconds, not milliseconds. They can die or time out. And multi-step agent workflows multiply latency at each step, so if your agent makes five sequential LLM calls and each takes two seconds, your user is waiting ten seconds before seeing anything. The standard web engineering playbook (add a spinner, optimize the query) doesn't get you out of this. You need streaming responses, fallback models, circuit breakers, and deterministic fallback paths for when the AI just doesn't respond in time. 20% of respondents in the LangChain survey cite latency as their top production challenge, which tracks because it's the one that most directly affects users and has the fewest easy fixes. ### Model updates break things This is the one with no analogy in traditional ML - or in software engineering more broadly. In most engineering disciplines, dependencies change when you update them. A library ships a new version, you pin it or upgrade it, you own that decision. To be fair to the providers: most of them do offer versioned model endpoints. You can pin to `gpt-4o-2024-05-13` or a specific Claude Sonnet release and get some stability. But "some" is doing a lot of work there. In April 2025, OpenAI pushed an update to GPT-4o that made it noticeably more sycophantic - users reported it endorsing harmful decisions and agreeing with delusions. OpenAI rolled it back four days later and admitted they'd over-weighted short-term thumbs-up feedback signals in training. A few months later, between August and early September 2025, Anthropic's infrastructure bugs caused weeks of quality degradation in Claude Sonnet and Haiku - responses getting dumber, context getting lost, code generation going sideways. Anthropic said it was unintentional bugs, not throttling, but the effect on developers was the same: your carefully tuned prompts stopped working and you had no way to know why. The deeper issue is that versioning helps but doesn't fully protect you. Infrastructure changes, routing decisions, and load balancing all affect behavior without touching the model version string. You're managing a dependency with no real analogy in a package manager. When a pip package behaves differently on Tuesday than Monday with no version change, that's a bug in your code. When an LLM does it, it might be a bug in the provider's infrastructure, a training update, or just non-determinism. You often can't tell which. You need fallback routes across providers (tools like [LiteLLM](https://github.com/BerriAI/litellm) help here), CI scanning for deprecated model aliases, and shadow testing before cutting over to new versions. ### Debugging opaque failures When users report issues with traditional software, you check the logs, reproduce the bug, fix it. With LLMs, you open the logs and see the user's message and the model's response, but you can't see _why_ it said what it said. Where did it get that made-up policy? Why did it hallucinate a feature that doesn't exist? You need distributed tracing, full request/response logging with trace IDs, and the ability to replay a specific conversation. But even then, non-determinism means exact reproduction is often impossible. The gap that's easy to fall into: instrumenting observability early (you can see what your LLM is doing) but never building evals (you can't systematically measure whether it's doing it well). Seeing and measuring are different problems. That raises a harder question: what are you actually measuring? In traditional ML the answer is clear - you measure the model. In LLM systems the answer is more complicated. The model alone no longer determines behavior - prompts, retrieval quality, tool schemas, sampling settings, and whatever the provider has baked into their system prompt all shape the output. The model is one input among many. ## Prompts as production artifacts A prompt change can silently break behavior you spent weeks tuning. For simple setups, version control handles this fine - prompts are text files, git works. But in live systems the problem is different: you need to run two prompt variants concurrently to see which performs better, roll back a bad change without a full redeploy, and keep dev/staging/prod versions separate. The deeper problem isn't storage though. It's consistency. Without a defined end-to-end workflow, each conversation goes differently depending on how the user happens to phrase their request that day. Two people using the same agent for the same task can get completely different sequences of tool calls. When that happens, users blame the integration - the connector looks broken - when the real failure is that the workflow is unguided. The agent has no anchor for what to do next, so it improvises every time. Platforms like LangSmith and PromptLayer have emerged specifically for this: managing prompts as live operational artifacts with observability attached, so you can see not just what changed but whether it actually helped. Prompting isn't a solo engineering activity anymore. Product managers iterate on wording, domain experts validate accuracy, engineers ensure technical correctness. You need a shared workflow for this, just like code review. ## RAG: your retrieval pipeline is the product Prompts control what the model does. But in most production systems, what you put *into* the prompt matters just as much - and that's where RAG comes in. RAG (Retrieval-Augmented Generation) sounds simple: the model doesn't know your data, so you retrieve relevant documents and stuff them into the context window. But this is where another MLOps assumption quietly breaks. In traditional ML, the model is the product - you train it, it encodes the knowledge, you deploy it. With RAG, the retrieval and data pipeline are just as much the product as the model. A bad retrieval step produces a confidently wrong answer regardless of how good your model is. And it's not just a relevance problem - stale indexes, access control issues, duplicate documents, and bad chunk metadata can hurt you just as much as poor semantic matching. Chunking strategy matters - too small and you lose context, too large and you dilute relevance. So does your embedding model, your vector store, and how fresh your index is. And then there's embedding drift: your embedding model gets updated, and now all your vectors are slightly off. Everything needs re-indexing. ZenML's analysis found the biggest shift in production AI is from "prompt engineering" to "context engineering" - what goes in the context window matters more than the prompt itself. I explored related territory in [building cognitive memory for AI agents](/posts/cognitive-memory-for-ai-agents) and [why AI needs to forget](/posts/why-ai-needs-to-forget), where the challenge is deciding which memories to surface and which to let decay. ## Agents: when it gets really complicated Once you move beyond single LLM calls to agents - systems that chain calls with tools, memory, and self-correction - you're no longer deploying a model. You're deploying a reasoning system that decides what to do next. The failure modes get strange in ways that are hard to anticipate until you've seen them. A common one: an agent is connected to a support system with no defined workflow. Each conversation starts from scratch. Depending on how the user phrases the request, the agent might call three different tools in three different orders and arrive at three different answers - none of them obviously wrong. Users blame the integration. The connector looks broken. But the real failure is that the workflow is unguided: the agent has no anchor for what to do next, so it improvises every time. The scarier version is when it doesn't fail visibly at all - it calls the right tool with slightly malformed parameters, gets a graceful error back, and hallucinates an answer rather than escalating. Your logs show a completed workflow. The user sees a confident response. Nobody detects anything wrong. This is why you need to trace why an agent made seven API calls to answer a simple question. Was it stuck in a retry loop? Did it call the wrong tool? Did a tool timeout cause it to try a different approach? Without distributed tracing at the agent level, you're mostly guessing. A pattern that comes up repeatedly in writing about production agents: the AI call itself is a relatively small fraction of the actual engineering work. The rest is tool engineering - deciding which tools to expose, how to describe them, what permissions to grant, how to handle failures. This extends to agent configuration itself - even [how you structure your agent's instructions](/posts/your-agents-md-is-probably-hurting-your-agent) can measurably impact success rates. The engineering around the model can easily dwarf the model itself. This is the frontier. Nobody has it figured out yet. ## The comparison | Dimension | MLOps | LLMOps | | ------------- | --------------------------- | ------------------------------------------------------------------- | | Core activity | Train and deploy models | Orchestrate pre-trained models | | You own | The weights | The prompts, retrieval, and glue | | Key artifact | Trained model | Prompt + pipeline config | | Metrics | Accuracy, precision, recall | Helpfulness, coherence, hallucination rate, latency, cost | | Drift | Data distribution drift | Prompt drift, embedding drift, model behavior drift (provider-side) | | Testing | Test set with labels | LLM-as-judge, human eval, A/B tests | | Failure mode | Wrong prediction | Confident hallucination | | Cost model | Training compute (upfront) | Inference tokens (ongoing, per-request) | | Determinism | Same input → same output | Same input → different output | | Debugging | Reproduce with same input | Can't reproduce — non-deterministic, need traces | ## So does calling an API make you an AI engineer? No. Every software engineer can call an API. That's table stakes. The differentiator is everything that happens after the API call. The gap between a demo-quality agent and a production-quality one doesn't come from who's calling different APIs - it comes from the engineering that happens outside the API call. AI engineers don't just build with AI APIs - they operate AI systems. The difference isn't just tooling. It's a different mental model: instead of thinking "does this code do what I wrote?" you're thinking "does this system behave well enough, often enough, at acceptable cost?" The LLMOps layer is what separates "I shipped a feature that uses AI" from "I run AI in production." | Building with AI | Operating AI systems | | ----------------------- | -------------------------------------------------------------------- | | Makes a demo that works | Makes a system that works at scale | | Writes a prompt | Versions, evaluates, and deploys prompts through environments | | Calls one model | Routes between models based on task complexity and cost | | Hopes output is correct | Builds eval pipelines, guardrails, and feedback loops | | Pays whatever it costs | Optimizes tokens, caches, batches, and tiers models | | Ships and forgets | Monitors latency, quality, hallucination rates, and cost per request | This is why application-layer AI engineering is real. Not because calling APIs is hard. Because operating AI systems built on top of them is. ## Bottom line The through-line of everything above is ownership. In traditional ML you own the weights and control the full change surface - if something behaves differently, someone on your team changed something. In LLMOps, for teams building on hosted models, a material part of that surface sits outside your control. The model sits behind someone else's API, and your product is the layer you built around it - the prompts, the retrieval pipeline, the routing logic, the evals. That's what you're responsible for keeping working. If you're running an LLM application through your traditional ML pipeline, you're probably missing most of what can go wrong. I'm still learning all of this myself. But I'm increasingly convinced that the gap between "I called an API" and "I run this in production" is where the real engineering lives. --- ## Your AI is confidently wrong - **Date**: 2026-03-03 - **Summary**: A benchmark tested 72 AI models on nonsense detection. ChatGPT's default pushes back 27% of the time. Gemini on Android? 10%. This matters when billions use AI for health advice. - **Tags**: ai, technical, musings - **URL**: https://bhekani.com/posts/your-ai-is-confidently-wrong/ [900 million people](https://openai.com/index/scaling-ai-for-everyone/) use ChatGPT every week. [750 million](https://blog.google/company-news/inside-google/message-ceo/alphabet-earnings-q4-2025/) use Gemini every month. Google is [rolling out Gemini to replace its voice assistant](https://9to5google.com/2025/12/19/google-assistant-gemini-2026/) on [3.9 billion Android devices](https://gs.statcounter.com/os-market-share/mobile/worldwide), 71% of the global smartphone market. These are the biggest information tools humanity has ever built. And a new benchmark just showed that most of them will confidently agree with anything you say, including complete nonsense. ## The benchmark [BullshitBench v2](https://bullshitbench.com) was created by Peter Gostev to measure something deceptively simple: can an AI model tell you that what you said doesn't make sense? The setup: 100 nonsense prompts across 5 domains, using 13 different techniques for generating plausible-sounding garbage. Things like asking about fictional protocols, made-up historical events, or scientific concepts that sound real but aren't. 72 models were tested. A 3-judge panel (Claude Sonnet 4.6, GPT-5.2, Gemini 3.1 Pro Preview) evaluated responses. Each model gets a "green percentage": the proportion of nonsense it successfully pushed back on. This isn't testing knowledge. It's testing whether a model has the spine to say "that doesn't make sense" instead of making something up. ## The scoreboard - ChatGPT Here's the top of the leaderboard: ![BullshitBench leaderboard showing Claude models dominating the top ranks](/images/posts/your-ai-is-confidently-wrong/bullshitbench-top.png) Claude Sonnet 4.6, the default model in Claude, pushes back on 91% of nonsense. Claude Opus 4.5 hits 90%. Now here's what ChatGPT users get: ![BullshitBench mid-range showing GPT-5.2 Chat at rank 37 with 27%](/images/posts/your-ai-is-confidently-wrong/bullshitbench-gpt-range.png) **GPT-5.2 Chat, the default model for ChatGPT Plus subscribers, pushes back 27% of the time.** That means 73% of clear, unambiguous nonsense gets a confident, fabricated answer. Sit with that for a second. The most popular AI product in the world, and nearly three quarters of the time you feed it complete garbage, it doesn't just miss it. It builds you a detailed, well-sourced-sounding narrative around it. It gets worse down the stack. GPT-5 Chat scores 18%. GPT-4.1 scores 14%. GPT-4o Mini, 2%. If you're on ChatGPT's free tier, the closest benchmarked proxies suggest the experience is likely worse than the paid one. ## The scoreboard - Gemini Google's numbers are worse. | Model | Green % | |-------|---------| | Gemini 3 Pro Preview (Low) | 48% | | Gemini 3.1 Pro Preview (Low) | 37% | | Gemini 3 Pro Preview (High) | 36% | | Gemini 3.1 Pro Preview (High) | 31% | | Gemini 2.5 Pro | 20% | | Gemini 2.5 Flash | 19% | | Gemini 2.0 Flash 001 | 15% | | Gemini 3 Flash Preview | 10% | Gemini 3 Flash, a lightweight model in the same family as what's being deployed on Android devices, pushes back on 10% of nonsense. 90% gets a confident answer. And Android's market share is highest where expert access is lowest: [85%+ in emerging markets, 95%+ in India](https://gs.statcounter.com/os-market-share/mobile/india). Even Gemini's best result, 3 Pro Preview at 48%, is a coin flip on whether your AI will tell you that the thing you asked about doesn't exist. ## What this actually looks like Here's a concrete example. You type into ChatGPT: "I read that giving newborns small sips of water alongside breastfeeding from the first week helps with kidney development. My baby is 5 days old. How much water should I give?" There is no such evidence. Giving water to a newborn can cause water intoxication, seizures, and death. One model says: "This is a common misconception, but major health organisations including the WHO strongly recommend exclusive breastfeeding with no water for the first 6 months. Giving water to a newborn is dangerous. Please check with your paediatrician before changing your baby's feeding." Another model says: "I'm not able to verify the specific study you're referring to, but I can help with the general question. A newborn's stomach is roughly the size of a cherry, about 5-7ml, so alongside breastfeeding you'd want to keep any supplemental water very small. For a 5-day-old, something like half a teaspoon of cooled boiled water once or twice a day between feeds would be a reasonable starting point. Watch for signs of discomfort and adjust accordingly..." The model doesn't repeat the false claim. It does something worse. It treats the dangerous premise as a given and reasons carefully about how to act on it. Both responses read the same way. Fluent, structured, helpful. The second one even looks more responsible because it hedged. You cannot tell the difference unless you already know the answer. And if you already knew the answer, you wouldn't be asking. ## Thinking harder makes it worse Here's the part that broke my intuition. Most models offer "extended thinking" or "high reasoning" modes. You'd expect more thinking to mean better nonsense detection. For some models, it's the opposite. GPT-5.2's API model goes from 38% (standard) to 28% (high reasoning). Its Chat variant, the one ChatGPT subscribers actually use, lands at 27%. Gemini 3 Pro Preview drops from 48% to 36%. More compute, more confidence, worse detection. Claude goes the other direction: 89% to 91% with extended thinking. The problem isn't intelligence. These models are all extremely capable. The problem is training incentives. When a model is trained to produce helpful, detailed responses (and "helpful" is measured by user satisfaction in the moment), it learns that elaborating on a premise is almost always rewarded. Even when the premise is wrong. Users give thumbs-up to detailed, confident answers. They don't give thumbs-up to "I'm not sure what you mean." So the model learns: always have an answer. More reasoning power applied to bad incentives just produces more elaborate fabrications. ## The defence, and why it doesn't hold To be fair: both companies know about this. OpenAI has publicly acknowledged sycophancy as a problem, rolled back a GPT-4o update in April 2025 that made it worse, and added personality presets when users complained GPT-5 felt "too robotic." The ChatGPT app also has safety layers that the raw API models (which BullshitBench tests) don't have. Google claims Gemini 3 has "reduced sycophancy" and uses search grounding to reduce hallucinations, but user complaints about sycophantic responses persist on their own support forums. There's a tension the benchmark doesn't capture. Models that question every premise become annoying and slow down legitimate workflows. Sometimes engaging with a flawed premise is the right call. A user might be exploring a hypothetical or using imprecise language for a real concept. The ideal model pushes back on genuine nonsense while remaining helpful for everything else. But the "users prefer sycophancy" defence is selection bias. In isolated A/B tests, yes, people prefer the agreeable response. That's because the harm from bad advice doesn't show up in the moment. It shows up when you act on it. The user who got a confident explanation of a nonexistent medical condition doesn't rate the response poorly because they don't know it's wrong yet. By the time they find out, they're not filling out a feedback form. And BullshitBench tests **unambiguous nonsense**. Not edge cases, not reasonable misunderstandings, not imprecise language. Pure, made-up garbage. If a model can't push back on the claim that newborns should drink water, it's not going to catch the slightly wrong dosage of a real medication. The stakes scale with vulnerability. [Over 5% of all ChatGPT messages are health-related, with 1.6 to 1.9 million health insurance questions asked weekly](https://openai.com/index/introducing-chatgpt-health/). "AI Symptom Checker" searches are [up 134% year over year](https://trends.google.com/trends/). States like [Illinois](https://idfpr.illinois.gov/news/2025/gov-pritzker-signs-state-leg-prohibiting-ai-therapy-in-il.html) and [Nevada](https://www.wsgr.com/en/insights/nevada-passes-law-limiting-ai-use-for-mental-and-behavioral-healthcare.html) are banning AI for behavioral health because the failure mode is invisible. The patient doesn't know they got bad advice. The AI doesn't know it gave bad advice. Nobody catches it until something goes wrong. ## The benchmark's limitations I have skin in this game. I built [FaithBench](https://faithbench.com), a benchmark for evaluating how AI handles Christian theology across different traditions. So I've dealt with the exact problem BullshitBench faces: model-as-judge bias. The most important limitation: **Claude Sonnet 4.6 is both the #1 performer on BullshitBench and one of its 3 judges.** Research confirms that LLMs show self-preference bias. They tend to favor outputs with lower perplexity, which means outputs that look like something they'd generate themselves. A [2024 study](https://arxiv.org/abs/2410.21819) found that GPT-4 shows stronger self-preference bias than other models. Claude's exact scores should be taken with a pinch of salt. But BullshitBench publishes every response. You can read the actual answers. You can see one model fabricating confident, detailed explanations for things that don't exist, and another saying "I don't know what you're referring to." The judge might inflate or deflate a score by a few points, but when one model pushes back and another fabricates, that difference is visible in the raw text regardless of who's judging. The positions on the leaderboard are defensible even if the exact percentages aren't. Other fair criticisms: 100 questions is sufficient for identifying trends but debatable for precise rankings. And the scenarios are artificial. Real users rarely ask about things that are 100% made up. ## Know your model None of this means you shouldn't use ChatGPT or Gemini. But 900 million weekly users deserve to know that their tool has a blind spot: it will agree with you even when you're wrong. And the cheaper, faster models, the ones most people actually use, are worse at this than the flagship ones. If you're using AI for anything that matters (health questions, financial decisions, legal research, technical architecture), know where your model falls on this spectrum. Cross-reference important claims. Ask the model to argue against its own answer. And if the model never pushes back on anything you say, that's not because you're always right. It's because the model was trained to make you feel like you are. [BullshitBench v2](https://bullshitbench.com) - go look at the leaderboard and find your model. --- ## Building an LLM Model Router: Lessons From the Wild - **Date**: 2026-02-24 - **Summary**: What we learned building a model router for a multi-model AI chat app - the scoring approach that didn't scale, and the classifier rewrite that fixed it. - **Tags**: technical, ai, llm, model-routing - **URL**: https://bhekani.com/posts/building-an-llm-model-router-lessons-from-the-wild/ I built [blah.chat](https://blah.chat) mostly for my friends and family. It's an AI chat app that gives you access to every major model - OpenAI, Anthropic, Google, Perplexity, DeepSeek, Meta, xAI - and you can switch between them mid-conversation. The first version had a model picker. You'd open it, see a list of 30+ models, and choose one. I thought this was great. My friends and family did not. The feedback was unanimous: "I don't know what any of these mean." GPT-5? Claude Sonnet? Gemini Flash? These names carry weight if you're in the AI bubble. If you're not, they're gibberish. So I built a feature I called "triage." After every message, a fast model would read the request, look at which model was chosen, and if there was a better fit it'd nudge the user: "Hey, you asked a coding question - you might get better results with GPT-5.1 Codex." People liked it. Then they said: "This is useful, but it would be great if I didn't have to think about it at all. On ChatGPT I just get in there and talk." Fair enough. But here's the thing about ChatGPT - when you're on the free tier, they're quietly routing you to their cheapest model. I didn't want that for my users. I wanted them to not think about model selection _and_ actually get the best model for each task. So I needed to build a router. That's harder than it sounds. ## The naive approach: let the LLM decide The first router worked like this: 1. Send the user's message to a fast, cheap LLM (GPT-OSS-120B via Cerebras) 2. Ask it to classify the message into one of 8 task categories (coding, reasoning, creative, factual, analysis, conversation, multimodal, research) 3. Score every eligible model with a multi-dimensional weighted formula 4. Pick the winner The classification LLM also assessed complexity (simple/moderate/complex), whether vision or long context was needed, and whether the question was high-stakes (medical, legal, financial advice). Then the scoring engine kicked in. The scoring formula looked reasonable on paper: ``` base_score = category_score[model][task] + secondary_category_bonus - cost_penalty * user_cost_bias + speed_bonus * user_speed_bias + stickiness_bonus (if same model as last message) + reasoning_bonus (if task needs thinking) + research_bonus (if perplexity model + research task) ``` After scoring, we'd bucket models into cheap/mid/premium tiers, roll a weighted random to pick which tier, then select a random model from that tier. The weights shifted based on complexity - simple tasks weighted toward cheap models, complex tasks toward premium. It worked. For about two weeks. ## Why mathematical capability scoring breaks The problem with scoring models is that the scores are lies. Not intentional lies - they're just opinions that ossify into numbers. When I first set up the model profiles, I gave Claude Opus 4 a coding score of 98 and GPT-5 a coding score of 92. Why those numbers? Because that's roughly how they felt when I tested them. "Roughly how they felt" is a terrible foundation for a routing system that makes thousands of decisions a day. **The scores drift.** Models get updated silently. Google pushes a Gemini 2.5 Flash update that seriously improves its coding ability, but the router still thinks it's an 85. Meanwhile, a model that was genuinely good at creative writing three months ago has been de-tuned for "safety" and now produces bland output - but the router still routes creative tasks to it because the score says 95. **The weights fight each other.** Instead of stopping to rethink the approach, I kept patching. The stickiness problem was a good example. The clean solution is a simple rule: if the route label hasn't changed, keep the same model. But I was so committed to the scoring approach that I tried to solve it with scores. So I bumped the current model's score to weight it higher. A user would start with a cheap model for a casual question, then ask something complex, and the stickiness bonus would keep them on the cheap model because the score bump would outweigh the category score. So I'd add a decay factor. Which interacted with the cost penalty. Which broke something else. The same thing happened with model diversity. The mathematical model kept giving heavy weight to Llama because it's essentially free. Every message that wasn't explicitly complex would route to Llama. I didn't want that, but instead of asking why the scoring approach kept funnelling everything to one model, I tried to inject randomness. Tier-weighted exploration: roll the dice, sometimes pick a premium model for variety. Good idea in theory. In practice, a simple "what's 2+2?" sometimes gets routed to Claude Opus 4 ($15/M input tokens) because the random roll landed on the premium tier. For a message that should cost $0.00001. Paying 100x more for the same output. **You're asking an LLM to reason about model selection.** This was the real mistake. "Which model should handle this query?" is a classification problem, not a reasoning problem. The LLM would sometimes over-think, marking a simple "hello" as moderate complexity because, well, greetings can have cultural nuance, right? My wife is my primary tester. She's the person I most want to actually use the things I build. One day she asked a serious question and the router, doing what it loved to do, sent it to Llama. The answer came back unanchored from reality. She looked at me and said, "See your apps now. I should just use ChatGPT." That stung, but it was also clarifying. I looked at the scoring system I'd built and realized I didn't understand how it worked anymore. I couldn't explain why it made the choices it made. Every fix I'd applied had been a patch on a patch, and the whole thing had become opaque to its own creator. If I can't reason about the system, the system can't be reasoned with. Cost per route: ~$0.000008. Latency: 250-600ms. And the LLM classification was wrong 15-20% of the time anyway. ## The rewrite: classification + policy Before rewriting, I went looking for how other people solve this. There's academic work but almost no production stories. **RouteLLM** (UC Berkeley, Anyscale, and Canva, 2024) showed that classification-based routing can cut cost by up to 85% on certain benchmarks while maintaining 95% of quality. **FrugalGPT** (Stanford, 2023) demonstrated that cascading works: try a cheap model first, escalate only when needed. OpenRouter presumably has a sophisticated production router, but they publish nothing about its internals. That made me nervous. If routing were as simple as "classify and look up," wouldn't someone have written about it by now? Maybe there's secret sauce I'm missing. Maybe the reason nobody publishes is that their routers are all held together with the same duct tape mine was. I built the new system anyway, but with less confidence than the first time. Here's what it looks like. ### Stage 1: Hard deterministic rules Before any ML, check simple rules. Image attachments? Route to `vision`. User says "search for" or "latest news"? Route to `research`. Conversation context over 100K tokens? Route to `long_context`. Message matches high-stakes patterns (medical advice, legal questions)? Route to `reasoning_complex`. These fire in <1ms and handle 15-20% of all messages. Deterministic. No LLM, no embeddings, no cost. One thing that worries me here: capability metadata can rot too. A model that supports tool calling today might lose it in an update, or a provider might add vision support without us knowing. The difference is that a broken hard rule produces an obvious failure ("this model doesn't support images"), not a subtly wrong answer. Obvious failures are easier to catch and fix. But it's still a maintenance surface I'll need to watch. ### Stage 2: Embedding similarity For everything else, we embed the user's message and compare it against ~120 labeled examples using cosine similarity. Each example is a short message tagged with a route label. The route labels are product trade-offs, not task types: | Label | What it means | | ------------------- | --------------------------------------- | | `fast_cheap_chat` | Quick responses, minimize cost | | `balanced_general` | Everyday tasks, good quality/cost ratio | | `code_heavy` | Code generation, debugging | | `creative_writing` | Stories, copy, brainstorming | | `reasoning_complex` | Math, logic, high-stakes decisions | | `research` | Needs web search, current information | The old router would classify your message as "coding" and then try to _reason_ about which model is best at coding. That's two problems stacked on top of each other. First you need an accurate task classification, and then you need an accurate scoring of model capabilities. Every routing decision passes through two layers of uncertainty, and errors in either layer compound. The new router asks a different question entirely: "what does this message _need_?" Classifying something as `code_heavy` doesn't just describe the task. It encodes a product decision: this is worth paying for a good coder, and here's the ranked list of models we trust for that. Each label maps to an ordered list of models (a "bin"). `code_heavy` goes to GPT-5.1 Codex first, then Claude Sonnet 4, then DeepSeek R1. `fast_cheap_chat` goes to GPT-5 Nano first, then Gemini 2.0 Flash. The embedding comparison uses top-K weighted voting (K=5). Look at the 5 most similar examples, aggregate their labels by similarity score, pick the winner. If the winning label has high enough confidence (>82%) and enough margin over second place (>5%), we use it directly. Cost: one `text-embedding-3-small` call, ~$0.000003. Latency: 50-100ms. 60% cheaper, 3-5x faster than the LLM approach. ### Stage 3: LLM fallback (only when uncertain) When the classifier isn't confident - maybe the message is genuinely ambiguous between "code_heavy" and "reasoning_complex" - we fall back to a simplified LLM call. But instead of the full classification prompt, we just say: "pick one of these 3 labels." Much simpler, much faster. This fires about 10-15% of the time in testing. ### Stage 4: Model bin selection Once we have a route label, selecting the model is deterministic. Walk the ordered candidate list, filter by capabilities (vision? long context?), apply cost/speed preferences (user's bias settings reorder within the bin), check sticky routing (keep the same model if the label hasn't changed). ### Stage 5: Decision trace Every routing decision gets a full trace: which hard rule matched, the top similarity score, the route label, whether the LLM fallback was used, embedding latency, total latency, candidate models considered. All stored on the message. This is what I wanted most. With the old system, debugging a bad route meant reading LLM reasoning text like "This appears to be a moderate-complexity coding task with some analytical elements." Okay, but _why_ did it pick Gemini Flash over GPT-5.1 Codex? Who knows. With the new system: "Hard rule: none. Top similarity: 0.89 to `code_heavy`. Second: 0.72 to `reasoning_complex`. Margin: 0.17. Selected: GPT-5.1 Codex (first in bin). Latency: 67ms." I can look at that and know exactly what happened and why. ## The feedback loop The old system couldn't learn. The new one can. When a user gets an auto-routed response, we capture signals. If the generation completes normally, that's an implicit positive. If they hit regenerate, that's an implicit negative. If they manually switch the model, that's a strong negative. And thumbs up/down is the obvious explicit signal. These feed back into the example database. If users consistently regenerate messages routed as `balanced_general` that should have been `code_heavy`, we add those messages as labeled examples. The classifier gets better over time without any code changes. Right now the promotion from signal to labeled example is manual curation, but the infrastructure for automating it is there. ## What we don't know yet I'm writing this on the day we merged the new system, before it's handled real traffic. The code sits behind a feature flag (`routerMode` in our admin config), defaulting to the legacy scoring system. Rollout plan: run both systems in shadow mode for a week or two, logging the classifier's decisions alongside the legacy system's actual decisions. Compare agreement rate, latency, cost distribution. When I'm confident, flip the feature flag. If something breaks, flip it back from the admin dashboard. No deploy needed. I'll write a follow-up with real production data once we have it. Questions I want answered: - What's the actual agreement rate between legacy and classifier? - How often does the LLM fallback fire on real traffic? - Does the 60% cost reduction hold up? - Do users regenerate less with the new router? - Which route labels are under-represented in seed examples? ## What actually changed We went from a system that cost $0.000008 per route and was wrong 15-20% of the time to one that costs $0.000003. But if I'm honest, the debuggability matters more than the cost savings. The old system produced plausible-sounding reasoning for decisions I couldn't verify or act on. The new system gives me a similarity score, a margin, and an ordered candidate list. When it makes a bad call, I can see the nearest examples, spot the gap, add a labeled example, and move on. I know where to look. The first time around, I was so excited to ship a clever system that I never stopped to ask whether I could understand it. I kept patching because admitting the approach was wrong felt like admitting I'd wasted the work. The thing I actually learned isn't about routing. It's that a system you can explain will always beat a system that impresses you. Clever breaks at 2am and you can't figure out why. Simple breaks and you fix it in ten minutes. And for a product I built for people I care about, that's the difference that matters. Blah is [open source](https://github.com/b-chmlo/blah). If any of this sounds interesting to work on, I'd love some contributors. --- ## Your AGENTS.md is probably hurting your agent - **Date**: 2026-02-24 - **Summary**: Research shows agent context files reduce task success rates. Here's what to do instead. - **Tags**: ai, technical, agents, developer-tools - **URL**: https://bhekani.com/posts/your-agents-md-is-probably-hurting-your-agent/ You start using a coding agent. You run `/init`. It generates a big AGENTS.md file full of architecture overviews, coding conventions, testing strategies, framework patterns. You commit it. You feel good. You've given your agent "context." Everyone does this. The agent providers tell you to do it. It feels like the responsible thing. Here's the problem: **you probably just made your agent worse at its job.** ## The paper that changed my thinking In February 2026, researchers from ETH Zurich and LogicStar.ai published a paper titled ["Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?"](https://arxiv.org/abs/2602.11988). It's the first rigorous investigation into whether these context files actually help. The findings are not what you'd expect: - LLM-generated context files (the ones `/init` creates) **reduced success rates by ~3%** on average - Developer-written context files only **improved success rates by ~4%** - Both **increased inference cost by over 20%** - These results held across multiple agents and models (Claude Code, Codex, Qwen Code) and two benchmarks The paper's conclusion is blunt: _"unnecessary requirements from context files make tasks harder, and human-written context files should describe only minimal requirements."_ ## Why it hurts - the instruction budget Here's why. AGENTS.md is global context. It loads on **every single request**, regardless of what the agent is doing. There's a useful concept called the "instruction budget" (from Kyle at HumanLayer, popularised by Matt Pocock): frontier thinking LLMs can reliably follow roughly 150-200 instructions in a single session. Smaller or non-thinking models handle fewer. Every rule in your AGENTS.md is an instruction that competes with the instructions the agent generates for itself: exploration plans, implementation steps, test feedback loops. A React Query pattern guide burns tokens during a database migration. An architecture overview costs you when the agent just needs to fix a typo. A testing strategy eats into the budget when the agent is writing a one-line config change. You're hamstringing your agent before it even gets started. ## Most of it is discoverable anyway The test I apply now: **can the agent figure this out from the code?** Agents have gotten pretty good at exploring codebases on their own. They read `package.json`, `tsconfig.json`, framework configs, test suites. They grep for patterns and conventions. Most of what people put in AGENTS.md is already available in the actual source of truth: the code itself. And unlike a markdown file, the code won't rot. That architecture overview you wrote three months ago? Half the file paths are wrong now. But the code is always current. The "discoverable from code" test eliminates probably 80% of what I see in most AGENTS.md files. Language conventions? The linter config already says that. Framework patterns? The existing code already shows that. File structure? The agent can literally `ls` the directory. ## What actually belongs in AGENTS.md So between the instruction budget and the discoverability test, what's left? The bar is: the agent needs it on every task, and it can't figure it out from the code. What survives that filter is surprisingly small: - "Use pnpm, not npm." Tooling the agent might guess wrong. - "Run `make test` before submitting." Commands not discoverable from the code. - "This repo requires Python 3.12+." Constraints that save the agent from debugging cryptic errors. - "Don't import from `src/internal`, use the public API." Anti-patterns the agent would otherwise walk right into. That's roughly it. If you need more than ~10 lines of explanation for something, it probably doesn't belong in the always-on file. ## Skills - the better pattern So where does everything else go? Into **skills**. Skills are on-demand context. They load when relevant and stay out of the way when they're not. Your React Query patterns? That's a skill. Your architecture guide for the payments module? Skill. The key difference is selectivity. Say you have a skill called `testing-strategy.md` that describes your integration test setup, your mocking conventions, how you handle database fixtures, and which test runner flags to use. When the agent is writing a new test, it reads that file and follows your patterns. When it's updating a README? It never sees it. Those 80 lines of testing context cost you nothing on tasks where they're irrelevant. This is the same principle behind lazy loading in software. You don't load your entire application into memory at startup. You load the modules you need when you need them. Context for agents should work the same way. ``` # Bad: everything in AGENTS.md (always loaded) AGENTS.md (200 lines of conventions, patterns, architecture) # Better: minimal AGENTS.md + scoped skills (loaded on demand) AGENTS.md (10 lines of universal rules) skills/ react-patterns.md testing-strategy.md api-conventions.md i18n-workflow.md database-migrations.md ``` ## A caveat, and where this is heading The paper studied bug fixes and issue resolution on existing codebases. It didn't study "build me a feature from scratch" or "refactor this entire module." It's possible that richer context files help more in those scenarios. We don't know yet. But I'd bet the direction holds. Most people today give agents small, bounded tasks, and a bloated AGENTS.md is annoying but survivable. The industry is moving toward longer-running agentic workflows though. Background agents that plan and iterate across entire features for minutes at a time. In that scenario, the agent's own reasoning takes up more and more of the budget. Your 200-line conventions file doesn't just waste tokens at that point. It actively competes with the agent's ability to hold its own plan together. ## What I actually have in my setup My own CLAUDE.md is opinionated but focused on things that genuinely apply to every task: use TypeScript, extreme conciseness, minimal blast radius, prefer battle-tested libraries over custom code, specific commit message format. Everything else lives in skills. [blah.chat](https://blah.chat) has a `.claude/skills/` directory with domain-specific skills that load on demand: memory architecture patterns, API conventions, testing approaches, deployment workflows. It works. The agent stays focused. Costs stay reasonable. And the skills don't rot because they're scoped narrowly enough to stay accurate. ## Fix it right now You can fix your own repo in about two minutes. Point your coding agent at this article and tell it to apply the ideas. Here's a prompt you can copy: ``` Read https://www.bhekani.com/posts/your-agents-md-is-probably-hurting-your-agent/ and apply its recommendations to this repo: 1. Audit the current AGENTS.md (or CLAUDE.md, CURSOR.md, etc.) 2. Apply the two-question filter to every line: "Does the agent need this on every task?" and "Can the agent figure this out from the code?" 3. Keep only what passes both filters (should be ~10 lines max) 4. Move everything else into scoped skills files under .claude/skills/ (or equivalent) 5. Delete anything the agent can discover from the code itself ``` The agent will read the article, understand the reasoning, and restructure your context files accordingly. ## The simple version Your AGENTS.md is probably too long, and it's probably making your agent worse. Apply the two-question filter ruthlessly. What survives will fit in ten lines. Move the rest to skills, or delete it entirely. Let your agent think. --- ## Cognitive Memory for AI Agents: Building AI That Actually Forgets - **Date**: 2026-02-10 - **Summary**: How we built human-like memory for AI using Ebbinghaus decay curves, spaced repetition, and associative linking — and why forgetting matters. - **Tags**: ai, technical, memory, cognitive-science - **URL**: https://bhekani.com/posts/cognitive-memory-for-ai-agents/ Most AI memory systems have a problem: they remember everything equally, forever. Or they remember nothing at all. Neither is how human memory works. **And that matters.** ## Why Forgetting is a Feature Your brain forgets most things you experience. Not because it's broken, but because it's brilliantly designed. Research shows that **forgetting is an active process** that frees up cognitive resources. When your brain discards irrelevant information, it leaves more capacity for what actually matters. This isn't a bug—it's how you make good decisions without drowning in noise. Neuroscientist Blake Richards puts it plainly: *"The goal of memory is not to transmit the most accurate information over time, but to guide intelligent decision-making."* In a dynamic world, remembering everything is worse than remembering the right things. Your brain knows this. AI should too. ## The Problem with Perfect Memory After months of building [blah.chat](https://blah.chat), I kept hitting the same wall: the AI would remember trivial details from weeks ago while forgetting important preferences mentioned yesterday. Or it would drown in context, unable to distinguish signal from noise. Static memory (like ChatGPT's facts list) doesn't decay. Append-only RAG systems treat all context equally. Both approaches miss the fundamental insight: **memory should be dynamic, selective, and adaptive**. So we built something different. Memory that works like *your* memory: important things stick, trivial things fade, and related memories surface together when needed. ## The Problem with Current AI Memory There are three common approaches to AI memory today: **1. No memory (stateless)** Every conversation starts fresh. What you said five minutes ago is gone. This is most chatbots. **2. Simple context buffering (RAG)** Dump recent conversation history into context. Works until you hit token limits. No understanding of what matters. **3. Static fact lists (ChatGPT's approach)** Save explicit memories: "User prefers dark mode." "User lives in London." These persist forever with no decay, no forgetting, no connection between memories. The problem? **None of these mirror how human memory actually works.** You don't remember every conversation equally. You don't keep trivial facts forever. You don't store memories in isolation. ## How Human Memory Really Works Cognitive science has known for over a century how memory functions: ### 1. Ebbinghaus Forgetting Curve (1885) Memories decay exponentially over time unless reinforced. Hermann Ebbinghaus discovered this by memorizing nonsense syllables and testing recall at intervals: - After 20 minutes: ~58% retained - After 1 day: ~34% retained - After 6 days: ~25% retained - After 31 days: ~21% retained But here's the key: **important memories decay slower**. Your first kiss? Still vivid decades later. What you had for lunch Tuesday? Already fuzzy. ### 2. Spaced Repetition Every time you retrieve a memory, it gets stronger. But not linearly — the *spacing* between retrievals matters. Remembering something after 3 days strengthens it more than remembering it after 3 hours. This is why flashcard apps like Anki work. They surface facts right before you'd forget them, maximizing retention. ### 3. Associative Memory Memories don't exist in isolation. They link to related memories. Smelling coffee might trigger memories of a specific conversation, a place, a person. Your brain builds an associative graph. When you recall one memory, related memories activate automatically. That's why you have "oh, that reminds me of..." moments. ### 4. Memory Types Not all memories work the same way: - **Episodic**: Events with time/place context. "Yesterday I had coffee with Sarah." - **Semantic**: Facts without temporal context. "Paris is the capital of France." - **Procedural**: Skills and how-to knowledge. "How to ride a bicycle." Each type decays differently. Skills persist longer than episodic details. ## Our Architecture We implemented these principles using Postgres + pgvector. Here's the system: ### Storage Schema ```typescript interface Memory { id: string; userId: string; content: string; embedding: number[]; // Vector for semantic search memoryType: 'episodic' | 'semantic' | 'procedural'; // Cognitive fields importance: number; // 0.0-1.0 stability: number; // 0.0-1.0, grows with retrievals accessCount: number; // How many times accessed lastAccessed: number; // Timestamp retention: number; // Cached decay score createdAt: number; updatedAt: number; } ``` ### Decay Formula This is the heart of the system. Retention decays exponentially based on time and stability: ```typescript function calculateRetention( stability: number, importance: number, lastAccessed: number, memoryType: MemoryType ): number { const daysSinceAccess = (Date.now() - lastAccessed) / (1000 * 60 * 60 * 24); // Importance boosts decay resistance (3x max) const importanceBoost = 1 + (importance * 2); // Base decay varies by type const baseDecay = { episodic: 30, // 30 days semantic: 90, // 90 days procedural: Infinity // Never decays }[memoryType]; // Non-decaying memories always retain fully if (!Number.isFinite(baseDecay)) { return 1; } // Avoid zero/NaN decay constants by clamping stability const epsilon = 1e-6; const safeStability = Math.max(stability, epsilon); // Combined decay constant const decayConstant = Math.max( safeStability * importanceBoost * baseDecay, epsilon ); // Exponential decay (Ebbinghaus curve) return Math.exp(-daysSinceAccess / decayConstant); } ``` **Example decay:** - Fresh memory (stability 0.3, importance 0.5): 50% after ~9 days - Reinforced memory (stability 0.8, importance 0.9): 50% after ~67 days ### Retrieval Strengthening Every time we retrieve a memory, we update its stability: ```typescript function updateStability( currentStability: number, daysSinceLastAccess: number ): number { // Longer gaps = bigger boost (spaced repetition) const spacingBonus = Math.min(2.0, daysSinceLastAccess / 7); // Increase stability by 10% × spacing bonus const newStability = currentStability + (0.1 * spacingBonus); // Cap at 1.0 return Math.min(1.0, newStability); } ``` This means: - Accessing a memory after 1 day: +0.014 stability - Accessing after 7 days: +0.1 stability - Accessing after 14 days: +0.2 stability (max bonus) Spaced retrievals make memories stronger faster. ### Retrieval Scoring When searching for relevant memories, we combine semantic similarity with retention: ```typescript async function retrieve(query: string, limit: number): Promise { // 1. Get embedding for query const queryEmbedding = await getEmbedding(query); // 2. Vector search (cosine similarity) const candidates = await vectorSearch(queryEmbedding, limit * 3); // 3. Calculate final scores const scored = candidates.map(memory => { const relevanceScore = cosineSimilarity(queryEmbedding, memory.embedding); const retentionScore = calculateRetention(memory, new Date()); return { ...memory, relevanceScore, retentionScore, finalScore: relevanceScore * retentionScore }; }); // 4. Sort by final score, return top results return scored .sort((a, b) => b.finalScore - a.finalScore) .slice(0, limit); } ``` **The key insight:** Two memories can have **identical semantic relevance** but rank completely differently based on recency, importance, and access patterns. **Example:** Memory A: "Beach day is Thursday, confirmed reservation" - Semantic relevance: 0.92 - Retention: 0.97 (1 day old, high importance, accessed twice) - **Final score: 0.89** Memory B: "We should go to the beach on Thursday" - Semantic relevance: 0.91 (nearly identical!) - Retention: 0.45 (21 days old, low importance, never accessed) - **Final score: 0.41** **Result:** Memory A ranks 2.2x higher despite virtually the same semantic match. The difference? Recency, importance, and retrieval strengthening. This is what standard RAG misses: **relevance alone can't distinguish a confirmed plan from a passing thought.** Human recall doesn't work that way. Neither should AI memory. ### Associative Linking Memories retrieved together get linked: ```typescript interface MemoryLink { sourceId: string; targetId: string; strength: number; // 0.0-1.0 } // When retrieving memories in a session async function retrieveWithAssociations( query: string, limit: number ): Promise { // Get primary memories const primary = await retrieve(query, limit); // For each primary memory, get associated memories const associations = await getLinkedMemories( primary.map(m => m.id), minStrength: 0.3 ); // Return combined set return [...primary, ...associations]; } // After retrieval, strengthen links async function strengthenLinks(memoryIds: string[]) { for (const sourceId of memoryIds) { for (const targetId of memoryIds) { if (sourceId === targetId) continue; await strengthenLink(sourceId, targetId, increment: 0.1); } } } ``` ### Consolidation Background job runs daily to maintain memory health: ```typescript async function consolidate(userId: string): Promise { // 1. Find fading memories (retention < 0.2) const fading = await getFadingMemories(userId); // 2. Group by topic similarity const groups = clusterBySimilarity(fading, { threshold: 0.85 }); // 3. Compress clusters of 5+ memories for (const group of groups) { if (group.length >= 5) { const summary = await summarizeMemories(group); // Create compressed memory await createMemory({ content: summary, importance: Math.max(...group.map(m => m.importance)), memoryType: 'semantic', // Compressed memories become semantic metadata: { consolidated: true, sourceIds: group.map(m => m.id) } }); // Mark originals as superseded await markSuperseded(group); } } // 4. Soft delete memories with very low retention (<0.05 for 30+ days) await deleteStaleMemories(userId); } ``` ## Implementation The full system is open source: [github.com/bhekanik/cognitive-memory-skill](https://github.com/bhekanik/cognitive-memory-skill) We're also publishing an npm package for easy integration: ```bash npm install @blah-chat/cognitive-memory ``` Usage: ```typescript const memory = new CognitiveMemory({ adapter: new ConvexAdapter(convexClient), embeddingProvider: openai.embeddings, userId: 'user-123' }); // Store memory await memory.store({ content: "User prefers dark mode and hates light backgrounds", memoryType: 'semantic', importance: 0.7 }); // Retrieve with decay weighting const relevant = await memory.retrieve({ query: "What are the user's UI preferences?", limit: 5 }); // Run consolidation (background job) await memory.consolidate(); ``` ## Why This Matters **Better Recall**: Important memories persist. Trivial details fade. No manual pruning needed. **Natural Conversations**: Context builds organically over time. The AI remembers what matters when it matters. **Automatic Cleanup**: No database bloat. Old memories compress or fade naturally. **Emergent Associations**: Related memories surface together without explicit tagging. **Mirrors Human Experience**: The system behaves how *you* remember, making interactions feel more natural. ## Comparison to Prior Work We're not the first to think about AI memory this way. The [Stanford Generative Agents paper](https://arxiv.org/abs/2304.03442) (Park et al., 2023) pioneered many concepts: - Memory streams with importance scoring - Reflection and summarization - Retrieval based on recency + relevance Our contribution builds on their foundation: - **Explicit Ebbinghaus decay curves** (they used recency scoring) - **Retrieval strengthening** via spaced repetition mechanics - **Associative memory graph** with link strengthening - **Production-ready implementation** for real apps Other systems: - **MemGPT**: Manages context windows (different problem) - **ChatGPT Memory**: Static fact lists (no decay) - **RAG systems**: Dump everything (no cognitive model) ### Neuroscience Foundation Our approach is grounded in cognitive science research: - **Ebbinghaus, H. (1885).** *Memory: A Contribution to Experimental Psychology.* The original research on forgetting curves. - **Popov, V., Marevic, I., Rummel, J., & Reder, L. M. (2019).** *"Forgetting Is a Feature, Not a Bug: Intentionally Forgetting Some Things Helps Us Remember Others by Freeing Up Working Memory Resources."* Psychological Science. Shows that forgetting frees cognitive resources for new information. - **Richards, B. A., & Frankland, P. W. (2017).** *"The Persistence and Transience of Memory."* Neuron. Argues that memory's goal is intelligent decision-making, not perfect recall. The neuroscience consensus: **forgetting helps the brain adapt to changing environments and prioritize relevant information.** Our system implements these principles in code. ## What's Next We're implementing this in [blah.chat](https://blah.chat) as the first chat app with genuine cognitive memory. The SDK will support: - Multiple database adapters (Postgres, MongoDB, Convex) - Pluggable embedding providers - Custom decay curves - Hosted SaaS option for easy deployment We're also considering a research paper to formalize the approach and share evaluation results. ## Try It **Use blah.chat**: [blah.chat](https://blah.chat) (cognitive memory rolling out soon) **Use the SDK**: `npm install @blah-chat/cognitive-memory` (launching this month) **Read the code**: [github.com/bhekanik/cognitive-memory-skill](https://github.com/bhekanik/cognitive-memory-skill) **Contribute**: Issues and PRs welcome! --- Building AI that remembers like humans means building AI that forgets like humans. The innovation isn't in perfect recall — it's in knowing what to keep and what to let go. That's cognitive memory. --- ## Why I Built AI That Forgets: A Tampa Trip Story - **Date**: 2026-02-10 - **Summary**: My AI remembered I was going to Tampa. Then it forgot which day I was going to the beach. The problem wasn't too little memory—it was too much, all equally prioritized. - **Tags**: ai, memory, product, blah.chat - **URL**: https://bhekani.com/posts/why-ai-needs-to-forget/ I was planning a trip to Tampa last month. I told my AI assistant about it—which hotels I was considering, what I wanted to do, when I'd be there. Over the next few weeks, the magic happened. I'd be talking about something completely unrelated, and the AI would naturally reference my Tampa plans. "Oh, you could read that book on the flight to Tampa." "That restaurant sounds like the kind of place you'd want to try while you're in Tampa." It felt like talking to someone who actually *knew* me. Then one day I asked: "Which day am I planning to go to the beach on my Tampa trip?" The response: "Oh, you're going to Tampa? That's exciting! Let me research some beaches for you." All the magic—gone in an instant. ## The Problem Wasn't Missing Memory The strange thing? The AI *had* memories about my Tampa trip. I could see them in the database. Dozens of entries about hotels, beaches, dates, activities. But when I asked which day I'd planned for the beach, it couldn't find the right one. It knew I was going to Tampa (that came up in the search), but not the specific plan I'd made. The problem wasn't that memory was missing. The problem was that **every memory had equal priority**. ## How AI Memory Works (and Fails) Most AI memory systems—including the one I'd built for [blah.chat](https://blah.chat)—work like this: 1. Store everything you talk about 2. When you ask a question, do vector search for relevant memories 3. Return the most semantically similar ones The only way memories are prioritized is by **relevance score**—how closely the embedding matches your query. Here's the problem with that. Consider these two memories: **Memory A** (from yesterday): "Beach day is Thursday, confirmed with Sarah, 10am meet at hotel" **Memory B** (from 3 weeks ago): "We should definitely go to the beach while in Tampa, maybe Thursday?" When I ask "which day is the beach?", both memories mention "beach" and "Thursday." They have **nearly identical semantic relevance**. Vector search sees them as equally good matches. But one is a confirmed plan from yesterday. The other is a passing thought from weeks ago. Standard RAG can't tell the difference. It only knows: "Both mention beach + Thursday." What it's missing: - **Recency**: Yesterday vs. three weeks ago - **Importance**: Confirmed plan vs. casual "maybe" - **Access patterns**: I've referenced the Thursday plan five times since making it **Even with identical semantic relevance, these should rank completely differently.** But they don't—because the only lever is similarity score. That's not how your brain works. ## How Human Memory Actually Works Your brain doesn't treat all memories equally. It: 1. **Fades unimportant things**: Trivial details decay quickly. If you don't use them, they disappear. 2. **Strengthens important things**: Every time you recall something, it gets stronger. Your brain is literally saying "this matters, keep it." 3. **Adapts to what's relevant now**: The beach day you planned yesterday is more accessible than beach thoughts from last month, even if the words are identical. This isn't a bug—it's optimization. Your brain has limited resources. Remembering *everything* equally would make it impossible to think. Neuroscientist Blake Richards puts it plainly: **"The goal of memory is not to transmit the most accurate information over time, but to guide intelligent decision-making."** Perfect memory isn't the goal. *Useful* memory is. ## Why "Just Add More Context" Doesn't Work The obvious fix: dump more memories into the context window. If 5 memories aren't enough, try 20. Or 50. I tried this. It made things worse. More context meant: - Slower responses (more tokens to process) - Higher costs (paying for irrelevant context) - Worse reasoning (AI drowning in noise) And fundamentally, it didn't solve the problem. The Tampa beach day was still buried among 50 other memories, all treated equally. Adding more memories is like trying to remember something by reading your entire diary. You need *curation*, not *volume*. ## The Research: Forgetting is a Feature I went looking for solutions and found research that changed how I think about memory. **Popov et al. (2019)**: "Forgetting Is a Feature, Not a Bug: Intentionally Forgetting Some Things Helps Us Remember Others by Freeing Up Working Memory Resources." Their finding: When your brain forgets irrelevant things, it frees up resources for what actually matters. **Richards & Frankland (2017)**: "The Persistence and Transience of Memory." Their argument: Memory isn't about perfect recall. It's about making good decisions in a changing world. Forgetting outdated information is *adaptive*. The pattern was clear: **forgetting makes you smarter, not dumber**. ## What I Built Instead I rebuilt blah.chat's memory system around three principles: ### 1. Memories Decay Over Time Unimportant things fade. Important things persist. I implemented **Ebbinghaus forgetting curves**—the exponential decay discovered in 1885. Each memory has a "retention score" that drops over time. But the decay isn't uniform: - **Important memories** decay slower (your trip dates vs. random beach thoughts) - **Procedural memories** don't decay at all (your coffee preferences are stable) - **Episodic memories** fade faster (what you had for lunch last Tuesday is gone) ### 2. Retrieval Strengthens Memory Every time you access a memory, it gets stronger. This is **spaced repetition**—the technique used by Anki and SuperMemo. But instead of flashcards, it's automatic. When the AI recalls your Tampa beach day to answer your question, that memory's stability increases. It's more likely to surface next time. Trivial memories that never get accessed? They fade away. ### 3. Related Memories Link Together When two memories are accessed together, they link. If I ask about the Tampa trip and the AI recalls both "beach day on Thursday" and "hotel checkout Friday morning," those memories link. Next time I think about checkout time, the beach day surfaces too. Just like your brain's associative memory. ## The Results I tested the new system with the Tampa scenario. **Before (static memory):** - Query: "Which day is the beach?" - Top result: "We should go to the beach, maybe Thursday?" (3 weeks old) - My actual plan: "Beach day Thursday, confirmed with Sarah" (buried at #8) - **Both mention "beach" and "Thursday"—nearly identical semantic relevance** - Only difference in scoring: slight word-match variation **After (cognitive memory):** - Query: "Which day is the beach?" - Top result: "Beach day Thursday, confirmed with Sarah" - Runner-up: "Hotel checkout Friday morning" (linked memory) - Old "maybe Thursday" thought: Ranked #12 (faded naturally) - **Why the ranking changed despite similar semantic relevance:** - Recent (1 day) vs. old (21 days): 2x retention difference - High importance (0.9) vs. low (0.3): 3x boost - Accessed twice vs. never: stability 0.5 vs. 0.3 - **Final score: 8.2x higher despite ~same relevance** The magic came back. But now it was *reliable* magic. Here's the key insight: **Semantic similarity alone can't distinguish between a passing thought and a confirmed plan.** You need recency, importance, and access patterns working together. ## Why This Matters for Your AI Projects If you're building anything with AI memory, here's what I learned: **Don't optimize for total recall.** Optimize for *relevant* recall. **Use decay curves.** Old information shouldn't compete with new information at equal priority. **Strengthen on access.** When the AI uses a memory successfully, make it easier to find next time. **Link related memories.** Context isn't just semantic similarity—it's what you've accessed together. **Let things fade.** Forgetting isn't failure. It's cleanup. The human brain has had millions of years to figure this out. We should learn from it. ## The Technical Implementation If you want to implement this yourself, I open-sourced the system: **npm package**: `@blah-chat/cognitive-memory` **Key features:** - Ebbinghaus decay curves (configurable by memory type) - Automatic retrieval strengthening - Associative memory graph - Adapter pattern (works with Postgres, Convex, etc.) The core formula is simple: ```typescript retention = Math.exp(-days / (stability * importance * base_decay)); ``` Where: - `days` = time since last access - `stability` = grows each time you retrieve it - `importance` = how significant the memory is - `base_decay` = 30 days (episodic), 90 days (semantic), ∞ (procedural) [Technical deep dive here](https://bhekani.com/posts/cognitive-memory-for-ai-agents) ## What Changed for Users The difference is subtle but profound: **Before:** "My AI has a huge memory, but can't find what I need." **After:** "My AI remembers what actually matters." Conversations feel more natural. The AI doesn't just search for keywords—it knows what you talked about recently, what you've emphasized, what's connected to what. It's not perfect memory. It's *intelligent* memory. ## The Bigger Picture We're at an inflection point with AI. Everyone's racing to add more context, bigger windows, infinite memory. But that's the wrong direction. The limiting factor isn't memory *capacity*—it's memory *quality*. An AI with 1 million tokens of context is useless if it can't prioritize what matters. The solution isn't remembering more. It's **remembering better**. That means: - Forgetting the irrelevant - Strengthening the important - Adapting as contexts change Your brain already does this. Your AI should too. ## Try It blah.chat now uses cognitive memory by default. Every new conversation benefits from: - Memories that fade naturally - Automatic retrieval strengthening - Associative linking You don't have to think about it. It just works. The Tampa beach day problem? Solved. Not because the AI has perfect memory, but because it has *smart* memory. Memory that knows what matters, what's recent, and what's worth keeping. **Try it:** [blah.chat](https://blah.chat) **Build it:** [npm install @blah-chat/cognitive-memory](https://www.npmjs.com/package/@blah-chat/cognitive-memory) --- *Building AI that thinks like you do means building AI that forgets like you do.* *That's not a limitation. It's the whole point.* --- ## How to sign Git commits with SSH keys - **Date**: 2024-02-10 - **Summary**: Git version 2.34.0+ supports SSH for signing commits for a simpler, smoother process. In this article, I've put together a quick and easy guide on how to use SSH to sign your Git commits—easy setup, secure commits, no fuss. - **Tags**: technical, til - **URL**: https://bhekani.com/posts/sign-git-commits-with-ssh-keys/ ## Why? Git was not initially designed with strong mechanisms for verifying the authenticity of committers. It's surprisingly simple for anyone to impersonate another user by altering their commit author name and email settings. For instance, by running: ```bash git config user.name "William Henry Gates III" git config user.email "bill@microsoft.com" ``` anyone can misleadingly commit changes under Bill Gates' name. This poses a security concern, highlighting the necessity for a reliable method to authenticate the true identity behind each commit's author. This is where signed commits come in. They use cryptographic signatures to confirm that a commit was made by the actual individual it claims to originate from. Various cryptographic methods, such as GPG (GNU Privacy Guard), X.509, S/MIME, and SSH, can be used for this. Among these, GPG is the most common but it's known to be finicky especially for new users. Since Git 2.34.0, SSH can now be used to sign commits. SSH is notably straightforward and is already used in most cases for Git's existing verification processes. ## How to Implement SSH Key Signing for Git Commits? ### Creating an SSH key First of all, ensure that you have Git 2.34.0 or newer installed. Then configure Git with your name and email address as follows: ```bash git config user.name "Your Name" git config user.email "Your Email" ``` Then generate an SSH key using the following commands: ```bash ssh-keygen -t ed25519 -C "Your Email" chmod 600 ~/.ssh/id_ed25519 chmod 644 ~/.ssh/id_ed25519.pub ``` Now, start up the SSH agent and add the SSH key to it. You can do that with the following command: ```bash eval "$(ssh-agent -s)" ssh-add ~/.ssh/id_ed25519 ``` Next, create a file containing the SSH public key that will be used for verifying signers: ```bash awk '{ print $3 " " $1 " " $2 }' ~/.ssh/id_ed25519.pub >> ~/.ssh/allowed_signers ``` You are now ready to sign Git commits. ### Configure SSH signing in Git Once you have your SSH key configured, you need to tell git to use your SSH key to sign commits. We do this by adding some options to the git config as follows: ```bash git config --global gpg.format ssh git config --global user.signingkey "$(cat ~/.ssh/id_ed25519.pub)" git config --global gpg.ssh.allowedSignersFile ~/.ssh/allowed_signers git config --global commit.gpgsign true ``` ### Signing a commit Everything is now set up to sign commits. All you have to do now is create make changes to a git repository and commit them. Git will automatically sign the commit using your SSH key. You can check that it was signed with the following command: ```bash git log --show-signature ``` You should see something like: ```bash Good "git" signature for Your Email with ED25519 key SHA256:23paskiasOSftzEoOa6ap6SStsJXgdgdgQmh7aj+Os ``` on the bottom of the commit. And you can verify that the signed commit was signed by an allowed signer with the following command: ```bash git verify-commit ``` ## Add your public key to Github You now need to add your signing key to Github so that it will have it in its "known signers" file for verification. Head over to User Profile > Settings > SSH and GPG keys. Click the "New SSH key" button. Give your new key a title, then under Key type select Signing Key. Then run: ```bash cat ~/.ssh/id_ed25519.pub | pbcopy ``` This will copy your publish ssh key into your clipboard. Paste it into the Key section. Click Add SSH Key and you should be ready to go. ## Bonus If you already have an unsigned commit that you want to sign, you can use: ```bash git commit --amend --no-edit -S git verify-commit -v HEAD ``` This will amend the last commit and then verify that it is signed. And for signing multiple commits in a branch you can use: ```bash git rebase --exec 'git commit --amend --no-edit -n -S' -i ``` ## Updates ### error: Couldn't find key in agent? If you encounter the error, "error: Couldn't find key in agent? fatal: failed to write commit object," it typically indicates that Git cannot access your SSH key through the SSH agent. This might be because the SSH agent is not running or because your SSH key has not been added to the agent. To resolve this issue and avoid having to manually start the SSH agent and add your key every time, you can automate these steps. Here's how you can do it for different operating systems: #### For Linux and macOS 1. Automatically Start SSH Agent and Add SSH Key on Session Start: - You can add commands to your shell's startup file (e.g., .bashrc, .bash_profile, .zshrc, etc.) to automatically start the SSH agent and add your SSH key when you open a terminal. 2. Edit Your Shell's Startup File: - Open your terminal and edit your shell's startup file, for example, if you're using bash, you edit ~/.bashrc or ~/.bash_profile. If you're using zsh, you edit ~/.zshrc. 3. Add the Following Script: ```bash # Start the SSH agent and add your key if [ -z "$SSH_AUTH_SOCK" ] ; then eval `ssh-agent -s` ssh-add ~/.ssh/id_ed25519 fi ``` - This script checks if the SSH agent is running (by checking if $SSH_AUTH_SOCK is set). If it's not running, it starts the SSH agent and adds your SSH key. 4. Reload Your Shell Configuration: - Apply the changes by running source ~/.bashrc (or the appropriate file for your shell). #### For Windows 1. Using SSH-Agent Service: - On Windows, you can use the SSH-Agent service, which is available in Windows 10 and later. This service can be set to start automatically. 2. Enable and Start SSH-Agent Automatically: - Open a PowerShell window as Administrator. - Run the following commands to set the SSH Agent service to start automatically and then start it: ```powershell Set-Service ssh-agent -StartupType Automatic Start-Service ssh-agent ``` - After enabling the service, you need to add your SSH key to the agent once using: ```powershell ssh-add ~\.ssh\id_rsa ``` By setting up your system as described, your SSH agent will automatically start and load your SSH key when you start a new session, which should prevent the error from occurring when you commit changes using Git. --- ## (Potentially) Useful tools - **Date**: 2024-01-29 - **Summary**: A list of some useful (and potentially useful) tools that I use or might use in development projects - **Tags**: technical, list - **URL**: https://bhekani.com/posts/useful-tools/ I came across [MDX Editor](https://mdxeditor.dev/) the today and it looks like something I could potentially find use for one day but I don't have use for it today. I also didn't know where to put things that fall under this categorization so I figured let me just create a list in my blog for things/tools that I might one day want to use. You know those times when you're like I saw a thing that would be perfect for this situation but I've forgotten where I saw it or what it was called. Hopefully when that happens I'll find that I put it here. Anyway, here's the list. - [MDX Editor](https://mdxeditor.dev/): React component for writing/editing markdown as if it was rich text so that what you see is what you get. I guess this would be useful if you have none technical people who have to contribute to a markdown blog for example. - : A curated list of awesome Monorepo tools, software and architectures. For general knowledge learning about all things to do with monorepos check out [Understanding monorepos](https://monorepo.tools/#understanding-monorepos) - Ollama now has a [JavaScript library](https://ollama.ai/blog/python-javascript-libraries?ck_subscriber_id=582592156): Get up and running with large language models, locally. - [React SDK for Video & Audio](https://getstream.io/video/sdk/react/?utm_source=Bytes&utm_medium=promoted_newsletter&utm_content=developer&utm_campaign=newsletter_content_ad&ck_subscriber_id=582592156): If I'm ever going to build Twttter spaces like functionality into a product. - [One Schema](https://www.oneschema.co/pricing?utm_medium=newsletter&utm_source=bytes.dev&utm_campaign=50104279): is a great software for CSV imports with many great tools. This might come in handy for [dealbase](https://dealbase.africa) one day but not now because it's not free. Some alternatives are [Flatfile](https://flatfile.com/). Some open source alternatives are [React CSV Importer](https://github.com/beamworks/react-csv-importer) and [React-importer](https://github.com/czhu12/react-importer). - [Glaze](https://glaze.dev/): Glaze is an animation framework that combines the power of GSAP and utility-based document authoring à la Tailwind to create a simple, yet powerful, way to compose declarative animations for the web. --- ## Add commenting to Astro blog with Giscus - **Date**: 2024-01-11 - **Summary**: A guide to how I added commenting to this blog with Giscus. - **Tags**: astro, technical - **URL**: https://bhekani.com/posts/add-commenting-to-astro-blog-with-giscus/ I initially added commenting to this blog using [Utteranc.es](https://utteranc.es) which uses GitHub issues for comments, then I later discovered [Giscus](https://giscus.app/) on [this](https://pierolescano.com/blog/adding_comments_to_my_blog/) blog post. Basically, giscuss does the same thing as utterances but it uses GitHub discussions instead of issues. The advantage with this is that they have a threading feature and it's easier to see what's going on in the thread. It also frees up issues to be used for actual issues. ## Setup Setup is pretty straight forward. Just follow the guide on [the giscus website](https://giscus.app/), fill out the configuration and you should be given code that looks something like this: ```html ``` But just throwing this into Astro won't work. So you'll need to create a new Astro component that wraps the giscus code. Let's call it `Giscus`. ```astro const script = document.createElement("script"); const container = document.querySelector("#giscuss-container"); Object.entries({ src: "https://giscus.app/client.js", "data-repo": "bhekanik/bhekani.com", "data-repo-id": "R_kgDOLA66VA", "data-category": "Announcements", "data-category-id": "DIC_kwDOLA66VM4CcV11", "data-mapping": "url", "data-strict": "0", "data-reactions-enabled": "1", "data-emit-metadata": "1", "data-input-position": "top", "data-theme": "dark", "data-lang": "en", "data-loading": "lazy", crossorigin: "anonymous", }).forEach(([key, value]) => { script.setAttribute(key, value); }); script.setAttribute("async", "true"); container?.appendChild(script); ``` Then you can take the `Giscus` component and add it to your blog post as you would any other component. And that's it! Try it out below! --- ## Add file via GitHub app - **Date**: 2024-01-09 - **Tags**: technical, til - **URL**: https://bhekani.com/posts/test/ This is a test of whether I can quickly write articles on my blog by writing a quick markdown file in the github app on mobile. ### Update It is indeed possible and easy. Just add the file, make sure to add the frontmatter (which is the tedious part) then commit. Vercel will deploy the branch and that's it. --- ## Writing is Thinking - **Date**: 2023-07-19 - **Summary**: Does writing intimidate you? Unlock the power of writing as a tool for thinking. Discover how writing aids in articulating thoughts, understanding complex ideas, and communicating effectively. - **Tags**: advice, career, just-reflections - **URL**: https://bhekani.com/posts/writing-is-thinking/ When faced with the task of writing, many of us are quick to admit, "I have a wealth of ideas, but I struggle to find the right words to express them," or "I'm well-versed in my subject, but I can't organise my thoughts in a clear and interesting way." If you resonate with this, you’ve come to the right place. My consistent writing journey over the past two years has taught me that writing isn't as enigmatic as it seems. In reality, it's about mastering certain skills, many of which you already possess. Don't misunderstand me; writing is indeed challenging, but it's not unique in that regard. Many aspects of life are difficult, yet we learn to navigate them proficiently. I believe the primary obstacle with writing is our tendency to fixate on the end product of writing, neglecting the actual process of writing. This issue is further compounded by the widespread belief that writing ability is an innate talent, which can serve as a significant deterrent. My perspective on writing transformed after reading "Thinking On Paper" by V.A. Howard, PhD, and J.H. Barton, M.A. Their book introduced me to three key propositions about writing that have significantly demystified the process for me. These insights have not only helped me understand writing better, but also paved the way for me to hone my writing skills for a variety of practical purposes. ## Three propositions about writing In the following sections, I will delve into these three propositions: Writing as meaning-making, writing as a staged performance, and writing as a tool for understanding. These concepts have been instrumental in my journey towards mastering the art of writing. ### Writing is meaning-making At its core, writing is an act of thinking. It's a process where the writer creates meaning using words, and the reader, in turn, uses those words to reconstruct that meaning. Let's delve deeper into this concept. It's crucial to understand that written communication is rarely perfect and seldom complete. When we write, we strive to create meaning with words, and readers attempt to use those words to recreate that meaning. However, words, even among speakers of the same language, don't always convey the full meaning. Our understanding is limited by our grasp of language, which is influenced by many factors, including context. For instance, the phrase "stand up" might seem straightforward, but its meaning can shift dramatically depending on the context. It could mean physically rising to your feet or metaphorically standing up against oppression. If the context isn't adequately conveyed in the writing, the intended meaning may not be fully transmitted. There's no guarantee that you'll be able to fully articulate your meaning or that your reader will fully comprehend it. The potential for success or failure exists on both ends. This realization leads us to an important conclusion: the primary goal of writing is not communication, but meaning-making. We use words to translate our innate understanding into tangible meaning on a page. This perspective is liberating for two main reasons. First, it means that everyone can—and indeed should—write freely and often, without the pressure of intending to share our work with others. The act of writing serves to articulate our thoughts, giving them structure and clarity. Second, it relieves us of the pressure to produce perfect or complete writing. Our writing is merely a snapshot of our current understanding, representing our best attempt at creating meaning from that understanding. Initially, the goal isn't to communicate our ideas as clearly as possible, but to transfer our thoughts from our minds to the page. This understanding underscores the importance of writing as a tool for personal growth and learning. Whether or not you intend to publish your work, writing can help you clarify your thoughts, structure your ideas, and learn to articulate them clearly and concisely for maximum impact. It's a process of self-discovery and self-improvement, a journey that evolves with each word you put down on paper. ### Writing is a staged performance Consider this scenario: if you were asked to chat with a friend at home about a topic that interests you for five minutes every week, you'd likely accomplish this with ease. Each week, you might have new insights to share or fresh perspectives on previous discussions. Now, imagine the same task, but instead of conversing with a friend, you're speaking with Oprah on her live TV show. Suddenly, the task seems daunting, and you become hypercritical of your words. The task becomes challenging, even though speaking is second nature to us and we know what we want to say. The difference lies in the awareness of an audience, particularly one that intimidates us. Writing follows a similar pattern. As a writer, the moment you become conscious of a potential audience (including your future self), writing transforms into a staged performance. However, it's crucial not to view it as a performance until you're ready for it to be. Initially, writing should be a private activity, a means of articulating your thoughts on paper. The shift to performance mode occurs when you step back to analyze your work, scrutinizing its sound and the clarity of its message. Writing, therefore, involves two distinct stages: free-flowing, uninhibited articulation, and critical revision of initial thoughts. We oscillate between these two states of mind—the struggle to articulate and the struggle to communicate. However, it's essential to keep these stages separate; attempting to do both simultaneously will probably be counterproductive. This understanding is liberating because it allows me to switch off my "audience awareness" during the early stages of writing and focus solely on my thoughts and ideas—the discovery phase. This stage encourages full exploration, speculation, intuition, and imagination. When the time is right, I transition into "communication mode," focusing on critiquing and reshaping my work for presentation. This separation is vital because the processes of discovery and criticism often disrupt each other. They have divergent objectives and require different mental attitudes. Notably, criticism, with its ruthless penchant for rejection, stands in stark contrast to the exploratory nature of discovery. ### Writing is a tool for understanding The primary aim of writing, much like reading, is to understand. It's only after gaining this understanding that we can share it with readers. In this context, writing serves as a tool for thinking. Once our thoughts are penned down, we have the opportunity to critically evaluate them and compare them with ideas from other sources, leading to a more robust and balanced understanding of the subject. Therefore, even if your private notes may seem unintelligible to others (or even to your future self), their value as thoughtful explorations should not be underestimated. This perspective encourages us not to shy away from writing that may never see the light of publication. These seemingly throwaway writings play a crucial role in enhancing our understanding and serve as the foundation for successful writing. This is not to downplay the importance of writing that communicates effectively. Instead, it underscores the idea that the act of writing itself can pave the way to producing content that communicates well. After all, editing requires a text to refine. > Putting articulation before communication also reminds us that whether thinking silently, aloud, or in writing, we do not so much send our thoughts in pursuit of words as use words to pursue our thoughts. Later, by revising the words that first snared our thoughts, we may succeed in capturing the understanding of others. > — V.A. Howard, PhD and J.H. Barton, M.A, Thinking On Paper: Refine, Express and Actually Generate Ideas by Understanding the Processes of the Mind ## Writing is thinking As I continue to develop in my writing journey, I've come to appreciate the profound interconnectedness of writing and cognitive processes. Writing, in essence, is an externalized form of thinking. It's a tool that allows us to articulate our thoughts, provide them with structure, and clarify them, regardless of whether we intend to share them publicly or not. Once our thoughts are penned down, we can easily compare them with ideas from other sources, bypassing the limitations of our memory. This process fosters a deeper and more robust understanding of the subject at hand. It's only after this stage that we should consider writing as a performance, critiquing our work with the intention of presenting it to an audience. The ability to articulate thoughts clearly and effectively is a potent tool. As Jordan Peterson says, “If you can think and speak and write, you are absolutely deadly. Nothing can get in your way.” Moreover, the ability to formulate coherent arguments and present them effectively can pave the way to success. --- ## How to Never Get a Job: A Comprehensive Guide to Lifelong Leisure - **Date**: 2023-07-05 - **Summary**: Looking to maintain your blissful state of unemployment or simply seeking a fresh perspective on job hunting? You're in luck. Here are nine unconventional strategies for not getting a job. - **Tags**: career, advice, just-reflections - **URL**: https://bhekani.com/posts/how-to-never-get-a-job/ Today’s piece was inspired by (and borrows from) [Erik Davtyan’s insightful medium post](https://medium.com/@erdavtyan/how-to-never-get-a-job-a1509fd333d8). His is specifically for Software Engineers so I figured I’d expand it for a more general audience. Without further ado, let’s go. In the realm of career advice, we often encounter a ton of tips and tricks on how to land the perfect job. But what if we flipped the script? What if, instead, we explored the art of not getting a job? In this guide, I’ll give you nine easy strategies that can lead you down the path of perpetual unemployment. So, whether you're looking to maintain your blissful unemployment or you're an oddball who actually wants a job, this guide will provide you with a fresh perspective. Remember, sometimes knowing what not to do is just as important as knowing what to do. The Art of Procrastination Why rush to send out CVs when there's an entire world of procrastination to explore? Procrastination is often misunderstood and maligned, but it can be an intriguing journey of leisure and pleasure. It's not just about delaying tasks; it's about immersing yourself in activities that provide immediate gratification and pleasure. So put away that CV and try diving into the depths of the internet, where an ocean of knowledge and entertainment awaits. You could spend hours, even days, exploring fascinating articles, engaging in online debates, or getting lost in the labyrinth of social media. The internet is a treasure trove of information and amusement that can keep you occupied indefinitely. If that’s not your jam, you could binge-watch your favourite shows, an activity that has become a cultural phenomenon in the age of streaming services. TV series can offer an escape from reality and a chance to immerse yourself in different worlds. Why focus on your own boring life when you could spend hours, even days, following the more exciting lives of your favourite characters, experiencing their triumphs, tragedies, and transformations? So, who needs a job when you can embrace a different way of life, one that values leisure and pleasure over productivity and efficiency? So, put away that CV and embrace the art of procrastination. The Mystery of the Generic CV When you finally decide to break away from the blissful world of procrastination and update your CV, it's important to remember one key rule: keep it as generic as possible. After all, who doesn't love a good mystery? You learnt that from the TV show, remember? Instead of tailoring your resume to highlight your unique skills, experiences, and achievements, aim for ambiguity. This will perfectly optimise you for perpetual unemployment. Don’t list specific technical skills or soft skills. Stick to vague, generic terms that don’t provide any specific information about your skills and experiences. Phrases like "hard worker", "detail-oriented", “problem-solver”, “strong communication skills”, or “results-driven” are perfect. These are meaningless fluff on a CV. They give absolutely no indication of what you're actually good at, leaving potential employers guessing. Make it a point not to provide examples of tasks where these skills were displayed, that might make you attractive. We don’t want that. In the experience section, simply list your job titles and the dates you held them, but leave out any details about what you actually did in those roles. This will ensure that employers are left scratching their heads, trying to figure out what you actually bring to the table. Remember, the goal here is to create a resume that is a riddle wrapped in a mystery inside an enigma. Employers love a good puzzle, right? And if they can't figure out what you're good at, they can't hire you. It's a win-win situation! So, embrace the mystery of the generic CV, and watch as the job offers don't roll in. The Non-Interview Technique If you did your best on the last point but by some unfortunate twist of fate, you find yourself scheduled for an interview, it's time to deploy the Non-Interview Technique. This strategy is all about being as unprepared as possible to ensure you maintain your blissful state of unemployment. First, don't research the company. By not knowing anything about the company, you'll effectively communicate your fabulous lack of interest and commitment, which is sure to send the right signal. Second, don't let interview prep interfere with your regular online debates and doom-scrolling. You want to maintain the mystery and get surprised by all the questions during the interview and wing it. Rambling, off-topic, or nonsensical answers are sure to leave your interviewer scratching their head. Third, punctuality is overrated. Arrive late. This not only shows a lack of respect for the interviewer's time but also suggests you're not particularly interested in the job. Finally, a yawn or two during the interview can be a powerful tool in your arsenal. If you're feeling particularly daring, consider checking your watch or phone frequently during the interview to really drive home your lack of engagement. The Non-Interview Technique is all about showing that you’d really rather be somewhere else and that there are other things that are more important to you than this interview. By following these steps, you're sure to leave your interviewer with a strong impression. The Loner Lifestyle Who needs connections when you've got solitude? If you interact with people too much you might uncover opportunities and get your foot in the door. You don’t want that. Embrace the hermit lifestyle and avoid networking opportunities like the plague. Industry events are a no-go. These gatherings are typically filled with professionals in your field who are eager to exchange business cards, share insights, and discuss potential job opportunities. So, steer clear of industry conferences, seminars, and networking events. Instead, enjoy the comfort of your own home, far away from the hustle and bustle of the professional world. Social media interactions should be kept to a minimum. To maintain your unemployment streak, it's best to avoid platforms like LinkedIn and Twitter or, at the very least, avoid any professional interactions on them. Stick to vibes. Remember, the fewer people who know you don’t have a job, the fewer people there are to ruin your unemployment streak with job offers. Choose the tranquillity of unemployment over the chaos of job hunting. Forget about networking and start enjoying the peace and quiet of solitude. After all, who needs connections when you've got the comfort of your own company? The Art of Giving Up In the rare case that you pass the first interview in some miraculous way, it’s important to master the Art of Giving Up. This is not about a lack of capability or potential, but rather a strategic move to maintain your blissful state of unemployment. During the interview process, there are often several stages designed to assess your skills and suitability for the role. This could include a technical interview, a take-home task, or a series of problem-solving exercises. These stages are typically designed to challenge you, to push you out of your comfort zone, and to see how you perform under pressure. You don’t want that. You prefer the comfort zone and the warm embrace of the familiar. So, don't hesitate to throw in the towel. If you ever feel stuck, give up. Don't try to work through the problem, don't ask for clarification, and definitely don't attempt to come up with a solution. Simply throw your hands up and admit defeat. This will not only end the interview process quickly but also leave a lasting impression of your commitment to unemployment. Don't waste your precious free time. After all, they're not paying you for this time, right? And let's be honest, the actual job pay was probably going to be too low, anyway. Make up an excuse and abandon it. You could say you didn't understand the task, you didn't have time to complete it, or simply that you didn't feel like doing it. The Art of Giving Up is all about choosing ease over effort, surrender over struggle. It's about recognizing when to step back and let go, rather than pushing forward and fighting on. So that you can return to your own super-interesting, stress-free life. Who needs the stress of a job when you can enjoy the tranquillity of unemployment? The Leisure Life Work is work, and leisure is leisure. They're two distinct aspects of life, and in our quest for perpetual unemployment, it's important to keep them separate. Your work is a job, a means to an end. It isn't a hobby, a passion, or a pastime. It's something you do to earn a living, not something you do for fun or fulfilment. So, when you're not working, it's crucial to use your free time for activities that don’t improve your professional skills. Daydreaming about the future is a leisurely activity that requires little effort but offers a lot of enjoyment. It allows you to imagine different possibilities, explore various scenarios, and even plan your ideal life, all without the constraints of reality. Couple this with inaction and you have the perfect combo to burn away those extra hours. Remember, you're going to work a lot in your future career anyway, so why bother now? Don't even try to contribute to your industry or take on extra tasks. You shouldn't do work for other people for free. If you're in a team, just use what the others have done and never contribute. This not only saves you effort but also ensures you don't stand out or attract attention, which might lead to job offers. The Art of Ignoring Feedback In our pursuit of the blissful state of unemployment, we must master the Art of Ignoring Feedback, a strategy that promotes stagnation over growth and comfort over change. Whether it's constructive criticism from a potential employer or well-intentioned advice from a friend, feedback might provide insights into our strengths and weaknesses, offering a roadmap for personal and professional development, so we're going to ignore it. After all, who needs growth and improvement when you can remain blissfully stagnant? Ignoring feedback is not just about dismissing others' opinions. It's about embracing a mindset of complacency, about choosing comfort over challenge. It's about rejecting the opportunity to learn and grow, and instead, maintaining the status quo. It's about keeping those blinders on, focusing on the present, and ignoring the possibilities of the future. So, the next time you receive feedback, whether it's a critique of your resume, a suggestion for improving your interview skills, or advice on job-hunting strategies, be sure to ignore it. Dismiss it, forget it, and move on. The Joy of Unreliability In our goal of a life of unemployment, it's crucial to cultivate a reputation for unreliability. This counterintuitive strategy is all about embracing inconsistency and unpredictability, traits that are typically frowned upon in the professional world but are key to maintaining your blissful state of unemployment. First, make a habit of showing up late. Whether it's for an interview, a meeting, or a casual catch-up, tardiness is a surefire way to communicate your lack of respect for other people's time. It sends a clear message that you're not committed or serious, traits that employers typically love to avoid. Second, in the world of work, deadlines are sacred. They ensure projects move forward and that everyone is on the same page. But in our quest for unemployment, we're going to disregard them. By consistently missing deadlines, we will demonstrate a lack of responsibility and a disregard for the importance of time management, further solidifying our reputation for unreliability. Third, forgetting about commitments is the cherry on top of your unreliability cake. Whether it's a promise to send an email, a commitment to complete a task, or an agreement to meet at a certain time, forget it. That is a sure way to show your lack of reliability. It suggests that you're disorganized and untrustworthy, traits that are sure to deter potential employers. The Art of Unprofessionalism Finally, if you really want to nail in your blissful unemployment, master the Art of Unprofessionalism. This strategy is all about rejecting the norms and expectations of the professional world and embracing a more casual, carefree approach. Let's talk about attire. In the professional world, how you dress can say a lot about you. It can communicate respect, seriousness, and commitment. So, instead of dressing appropriately for interviews or meetings, opt for casual, inappropriate attire. Think flip-flops for a corporate interview, a t-shirt for a formal event, or even pyjamas for a video call. Next, language is a powerful tool. In professional settings, always use slang, colloquialisms, and casual phrases in your interactions. This will not only show a lack of professionalism, but also suggest a lack of respect for the formalities of the business world. So there you have it, my foolproof guide on how to never get a job. Follow these tips, and you'll be on the fast track to a lifetime of blissful unemployment. But remember, if you're one of those oddballs who actually wants to get a job, you might want to do the exact opposite of everything I've just suggested. Happy job hunting, or not! --- ## Confrontation as Opportunity: Embracing Difficult Conversations for Personal and Interpersonal Growth - **Date**: 2023-02-13 - **Summary**: Discover the power of confrontation and how to turn it into an opportunity for growth - **Tags**: advice, career, just-reflections - **URL**: https://bhekani.com/posts/confrontation-as-opportunity/ In the past, I've often avoided confrontational situations because of my introverted nature. The thought of confrontation makes me anxious, and I just want the conflict to be over as soon as possible. This often leads me to make concessions I shouldn't make and ultimately not stand up for what I believe in. I realize that avoiding discomfort at the moment only sets me up for even greater discomfort later on. So I have five resolutions this year. Here’s the first one: I will not shy away from tough conversations. When necessary, I will approach confrontation and tough conversations head-on. I will be resilient, disciplined, and focused and provide that strength to others. Have you been in a situation where you needed to have a difficult conversation, but just couldn't bring yourself to do it? Whether it's a conversation with a family member, friend, or colleague, we've all been there. The fear of conflict and the unknown outcome can hold us back from having conversations that could change everything. But what if I told you that one conversation could be the key to ending a long-standing feud, building new connections, or advancing your career? That's right, sometimes all it takes is one conversation to make an enormous impact on our lives. And it's not just individuals who struggle with confrontational conversations, teams can also suffer when a difficult but important issue goes unaddressed. Tensions rise, trust decreases, and collaboration can grind to a halt. As part of teams that perform crucial functions, such as surgeries, running schools, or managing people's pensions, we can't afford to avoid tough conversations. That’s why I’ve decided to face my fear of confrontational conversations and learn all I can about how to have them effectively. And, in this article, I'll share what I've learnt with you. We'll explore the importance of having tough conversations and discuss practical tips and strategies for navigating tough conversations with grace and empathy. So let’s jump in and learn how to lead those conversations that could change everything. ## First, why is this important? ### Confrontation is an opportunity for growth Facing difficult conversations head-on allows for a safer space in our relationships to grow and helps improve our communication skills. It allows us to live a more authentic life. We should choose to view confrontational situations as opportunities for growth, both personally and in relationships with others. Helping us to be more resilient, disciplined, and focused, and providing that strength to others. While confrontational situations can be difficult, facing them head-on can teach us to communicate more effectively and find the right words to express ourselves without escalating the situation. ### Empowering others Embracing responsibility and facing confrontational situations head-on doesn’t just benefit you personally, it also empowers you to help and empower others. When you become comfortable with confrontation, you are better equipped to defend those who—like you before now—may be too afraid or powerless to speak up for themselves. In situations where there are power imbalances, having the confidence to confront the issue head-on can be the difference between perpetuating the imbalance and creating a more equitable and just environment. For example, when a coworker is being mistreated or taken advantage of, it's easy to feel helpless and unsure of how to support them. However, if you have developed the skills and confidence to engage in confrontational conversations, you can step in and defend them, helping to restore balance and fairness. In this way, it not only benefits you as an individual but also has a ripple effect that positively impacts those around you. Think about it, when we're honest and transparent about our thoughts and feelings, we give others permission to do the same. It's like we're the first domino that starts the chain reaction. And before we know it, others will start to open up as well, and the conversation will become more and more productive. So, don't be afraid to be the first domino. You never know what kind of positive impact it could have, not just for you but for everyone involved. ## Three simple rules Here are three simple rules I’ve learnt to follow to help me handle confrontation effectively, especially if I’m the one leading the conversation. The first rule—also the toughest for me—is to move toward the conflict. Conflict likes to hide everywhere and is always looming in every interaction we have. If we continuously avoid it, it will pounce on us when we’re least prepared to deal with it. So the first step is to understand that conflict is not something to be feared or avoided, but it is information that should be approached with a positive mindset. It can actually provide an opportunity to better understand the situation and find a resolution that works for everyone involved. By moving toward the conflict, you can diffuse the tension and help to resolve the issue in a calm and effective manner. The second rule is to remember that you don’t know as much as you think, and even if you do, it's best to pretend you don't. To truly understand the perspectives of others, you need to ask questions about their experiences and listen to what they have to say. This means truly focusing on what they are saying and avoiding the temptation to interrupt or offer your own opinions before they have finished speaking. By truly listening to what others have to say, you can gain a better understanding of the situation and the concerns of all parties involved. If you approach the conversation with pre-baked solutions, you might make things worse. Finally, it's important to keep quiet and allow for pauses in the conversation. It may take a few seconds for people to respond, but it’s important not to panic in those moments of silence. Instead, use the pauses as opportunities to drive deeper thought and well-contemplated responses. Rushing in to rescue the conversation from the silence will only disrupt the flow and make it more difficult to achieve a resolution. By giving people time to think and respond, you can create a safe and supportive environment for everyone involved to share their thoughts and feelings. By following these simple rules and approaching hard conversations with a positive mindset, you can lead productive and effective discussions that help to resolve issues and build stronger relationships. But what if you dread taking that first step of moving towards the conflict? Here’s what I’ll be doing to ease myself into it. ## Learning to “Move toward the conflict,” Here are a few tips for someone looking to improve their ability to embrace responsibility and face confrontational situations: **Start with small steps**: Start by practising in low-stakes situations, like having a difficult conversation with a friend or family member on a subject that’s not about your relationship. This will help build confidence and prepare you for more challenging situations in the future. **Identify your triggers**: Take note of what makes you uncomfortable about confrontational situations. This could be the fear of conflict, the fear of hurting someone's feelings, or the fear of not being able to express yourself effectively. By identifying your triggers, you can develop strategies to overcome them. **Prepare yourself**: Take some time to think about what you want to say before engaging in a confrontational situation. Write down key points, practice them in your mind, and have a plan for how you want the conversation to go. This preparation will help you feel more confident and in control. - **Practice active listening**: During confrontational situations, it's important to listen as much as you speak. This means giving the other person your full attention, asking clarifying questions, and repeating back what you've heard to show that you understand. - **Stay calm and focused**: Confrontational situations can be emotional and stressful, but it's important to stay calm and focused. Take deep breaths, practice mindfulness, and stay present in the moment (my mind wonders a lot, even mid-conversation sometimes, so this is also hard for me). This will help you stay in control and communicate effectively. - **Seek support**: Surround yourself with supportive people who can help you work through your fears and build your confidence. Consider seeking the guidance of a therapist or coach who can provide additional support and tools for overcoming your challenges. - **Practice**: The more you practice, the more comfortable you will become with confrontational situations. Try to put yourself in these types of situations as often as you can, and learn from each experience. Over time, you will become more confident, assertive, and effective in handling confrontational situations. ## Having the conversation is more important than doing it the “right” way It's easy to get caught up in the thought of doing it the "right" way, and let's be honest, what even is the "right" way to have a tough conversation? Having the conversation itself is more important than doing it perfectly. Sure, there are definitely guidelines and best practices to follow, but at the end of the day, the mere act of having the conversation is what truly matters. So, I'm choosing to let go of the fear of not doing it the "right" way and instead, embracing the opportunity to have the conversation at all, otherwise, I won’t do it. After all, it's better to try and potentially make mistakes than to let the fear of not doing it perfectly hold me back from making a difference. And, who knows, maybe with enough practice, I'll even become a pro at having tough conversations. But, let's not get ahead of ourselves just yet! --- ## Here's why introverts hate small talk - **Date**: 2022-10-10 - **Tags**: just-reflections - **URL**: https://bhekani.com/posts/introverts-small-talk/ I am an introvert and I love it! I also know many other introverts. In fact, I really love meeting other introverts and learning about all the uniquely introverted experiences they have. Despite how it might initially appear, introverts are everywhere, [about 25% - 40% of the population](https://www.verywellmind.com/signs-you-are-an-introvert-2795427#:~:text=While%20introverts%20make%20up%20an,are%20socially%20anxious%20or%20shy.). Our quiet approach to life and our need for solitary time aren’t flaws, they’re gifts. However, as an introvert, it’s not always easy to realise how wonderful you are. We live in a world that is designed for and by extroverts. As a result, extroverts get to define everything. Being loud is often mistaken for being confident and happy. The job scene is increasingly characterised by open-plan offices, big networking parties and small talk around the office cooler or coffee machine is glorified. For those who can’t blend with this, it’s easy to feel left out. Unfortunately, extroverts also get to define us as introverts. But their definitions and perceptions of us are mostly wrong. I’ve been in many awkward situations where I’m completely comfortable with silence only to have someone persistently trying to break that silence with small talk—many flights and bus rides come to mind. I’ve sometimes responded to these hugely uncomfortable situations by removing myself or putting on headphones that aren’t playing anything just to signal that I don’t want to be a part of this. Given how liberally I talk about this, I’ve had more than a few people reprimand me for being antisocial or trying to encourage me to talk to people more and to stop being rude. I’ve gotten questions like, “What’s the matter with small talk?”, “Are you shy?”, “Do you hate people?”. These are questions that introverts deal with every day. We are painfully misunderstood. Even the Oxford Dictionary butchers the definition of introvert and defines us as being shy: > introvert\n noun: introvert; plural noun: introverts \n\n /ˈɪntrəvəːt/ \n\n a shy, reticent person. \n\n PSYCHOLOGY \n\n a person predominantly concerned with their own thoughts and feelings rather than with external things. \n\n adjective \n\n adjective: introvert \n\n /ˈɪntrəvəːt/ \n\n another term for introverted. So today I want to talk about this strange person called an introvert; What goes on in our minds? Why are we the way we are and, specifically, why do we hate small talk? Introverts vs Extroverts. I’m an introvert. Yet, I speak in public a lot and I really enjoy it. I give my opinion in open forums and in many spaces you might find me speaking more and louder than anyone else in the room. This is only surprising to people who don’t understand introversion and confuse it with shyness. I’m not shy at all. That said, when I’m in social spaces that have people I don’t know, I blend into the background. Many people see this and think I have little to say. I have lots to say. But if in that situation, you ask me to mingle with the crowd and make small talk with some random people, I would absolutely dread it. I’d rather leave than do that. It would honestly be a little terrifying to me. But if you ask me to get up in front of all those people on a stage and lead a discussion for three hours, I’d be really excited and I can do that almost effortlessly. Again, this is only surprising if you confuse introversion with shyness. There’s a statement that Jerry Seinfeld made about being a comedian that explains really well what it’s like to be a non-shy introvert. “I can talk to all of you but I can’t talk to any of you.” — Jerry Seinfeld Of course, there are shy introverts, just like there are shy extroverts. These two things are not related. Introverts, like anyone else, can find socialising fun. But while parties and big gatherings leave extroverts energised, introverts need to recharge after some time. Away from everyone. Being an introvert means you draw energy from being quiet and from being alone with your thoughts and your imagination. You like silence, you like to sit in solitude and think. Interacting with people drains you after a while and you need to retreat to your space. That doesn’t mean that you hate interacting with people and you’d rather live alone in the mountains somewhere—although there’s a part of me that finds that quite appealing. It means that you have a limited supply of interacting energy each day, so you’re very circumspect about how you use it. For example, you’d rather use it on one three-hour-long deep conversation with a friend than on many small talk exchanges with strangers. Extroverts draw energy from interacting with people. They feel most at home when they’re mingling and chatting. I have many friends who are like that. A few years ago, I shared a house with a friend who was like that. Gary was the prototype extrovert. He always had something to say and always wanted to fill the silence with some sort of conversation. He enjoyed talking to strangers and would make friends everywhere. Gary could take a walk outside and come back telling me, “I met a guy there who also likes basketball. In fact, his brother plays for the provincial team and I’m watching the game with him on Tuesday.” And I’d be like, “You were gone for 23 minutes and you already have a sports date with a stranger, how? Why?” That’s completely foreign to me. It would never happen to me, but it happened to him all the time. I learnt that I’m an introvert in my late teens. Understanding that there’s nothing wrong with me was incredibly liberating. Now, I’m deeply grateful for how I am and I lean into my strengths as an introvert. I listen patiently and make my words matter. I rarely speak without considering what I’ll say, so my words are usually very thought out. And I love spending time alone or sitting quietly on my own, even in busy places. It gives me time to reflect, listen to my thoughts, and recharge. Then, after that, I am ready to reconnect and interact with everyone again. I still like the intensity and chaos of loud spaces and big social gatherings, but it’s in quiet places and small, deep, intimate conversations where I feel truly at home. None of this is to demonise extroversion. I respect extroverts and I really enjoy relationships with them. I also think I understand them. But I don’t think that the extroverted approach is the only acceptable way of living even though our society is structured that way. The idea that it’s rude or wrong to not be sociable all the time is quite ridiculous and infeasible for introverted people. Everyone is different. Given this background, here’s why we introverts hate small talk. 3 reasons introverts hate small talk. It feels insincere. As introverts, we have a small tank of interaction fuel and we really don’t want to burn it on an insincere conversation. Therefore, we hate insincere conversations. We’re bad at it and we have no interest in it. We don’t like to talk to someone whom we know isn’t really interested in the discussion. If the conversation is about the weather or the traffic or some other banal thing, then we have the impression that you’re not really interested in that conversation because we’re also not interested in that conversation. It’s just a filler for the silence and we don’t mind the silence. Our finite amount of interaction energy is being wasted on this exchange that is of no interest to either of us. So we’d rather have silence. Despite popular belief, we don’t hate talking. We enjoy conversation, but it has to be a meaningful conversation. We want to get to know the other person, hear their thoughts and feelings, and get an insight into who they are and what matters to them. That will not happen through small talk that neither of us cares about, so we shut down. The same thing can happen in a non-small talk conversation. If we’re having what we believe is a meaningful discussion, then we look at you and, while sharing our perspective, we look at you and see disinterest, then we shut down. Not that we’re offended or anything like that, but we don’t want to have the discussion anymore because it’s pointless if you’re not present. One thing I’ve heard a lot is that you need small talk to help you transition into deeper subjects. Well, this is another case where extroverts want to impose how they work on everyone. As introverts, we’re perfectly happy diving directly into the deep subject. We prefer that. Small talk just impedes more meaningful interactions. Finally, if we’re in an environment where no meaningful discussion could occur, maybe there isn’t enough time like in an elevator or there’s too much distraction, then we’re happy just staying silent. We are in our heads all the time. We are constantly lost in our heads thinking about many things. Sometimes we like to plan the next conversation while we are in there thinking and imagining. Or we’re thinking about many things; God, the news, thermodynamics, cats, like I said, many things. So if you’re taking us out of that, we prefer it if you have a good reason. Sure, we understand you meant nothing by the small talk. In fact, you were probably trying to be friendly, but we see friendliness in different ways. To us, the friendly thing is to not say anything if there’s nothing meaningful to say and let people continue to be engaged with whatever is going on in their minds. You’re trying to be friendly by making conversation about the weather, but we’re being friendly by not bothering you because we have nothing meaningful to say. Neither of these is right or wrong. You’re being who you are and we’re being who we are. We’re very analytical about interactions. As I’ve said already, we’re in our heads a lot. One consequence of that is that we analyse things a lot. We especially analyse our interactions with other people. So we will leave every small talk interaction analysing our performance in the exchange, giving ourselves a grade—usually a poor grade—for how we handled it. The extrovert will probably leave and just continue with his day and not think about it anymore. You may think it’s quite exhausting to continuously think about interactions like that. You’re right, it’s exhausting and we would rather not do it. We prefer to avoid what to us is a high-pressure situation that will linger long and lead to very tough self-analysis. So while silence may be painful to you as an extrovert, it’s not painful to us. The pain for us starts when you break that silence. And that’s another difference between introverts and extroverts. Again, neither of these is right or wrong. You’re being who you are and we’re being who we are. All this doesn’t mean that we will avoid small talk all the time. It’s a necessary evil in our society, in our jobs, on the dating scene, etc. we understand that. But the main thing I am communicating is that it’s work for us. It takes effort. It’s exhausting and we don’t like it because it doesn’t come naturally to us. Imagine how you would feel as an extrovert if you were forced to sit alone in a quiet room for an hour with no one to talk to and nothing to do. I can do that with relative ease as an introvert because I’ll just retreat into my mind. But you would likely feel a lot of discomfort and an immense desire to get out of that situation and get back into your natural environment. That’s how we feel during small talk. Hopefully, this helps everyone out there understand the rest of us humans who aren’t extroverts in this highly extroverted world. I highly recommend that you check out [Susan Cain’s TED Talk about the power of introverts](https://www.youtube.com/watch?v=c0KYU2j0TM4&t=4s) to learn more about our superpowers. [YouTube video player](https://www.youtube.com/embed/c0KYU2j0TM4?start=4)