Voice agents for B2B, and an alternative data platform for market researchers, analysts, and investors.
The agent that answers the phone, and the operations layer behind it. Healthcare first: the agent is the front door, and the product is the work it triggers: intake, scheduling, insurance eligibility checked at the moment of booking, reminders, and the records that keep billing clean.
An alternative data platform for market researchers, analysts, and investors. Signal overlays that sit on top of price data, scenario testing to check how a thesis holds up, and one workspace to run both. On v12. A 200-member research community, 60 to 80 active testers per release.
I grew up in India and came here to find out what the other half of the world actually felt like. Penn State is where I landed. Nobody warns you about the part that sticks: you end up with two homes, and you are always a little homesick for whichever one you are not standing in.
Now I split my time between San Francisco and New York. I run on weekends. I try every cafe that opens near me, including the ones that are obviously going to be bad. I host small events around the neighborhood, mostly as an excuse to meet people who do things I have never thought about. Most of my good ideas started as someone else’s.
I like numbers more than is strictly healthy. I’ve built prediction models for markets and spent longer than I’ll admit looking for patterns in sports results. Cycle X and Three Axis both came out of that.
I work across tech, finance, and ops. I’ve spent a lot of time inside systems that break, and more time fixing them.












Case study · Prototype
Nobody asked me to do this. I built it unprompted as a product exercise for Canva, start to finish: the problem, the system, a working prototype and the ship plan.
A campaign is one message rendered into twenty formats. An Instagram post, a landing page hero, an email header, a printed flyer that somebody approved in a meeting nobody remembers. Design tools are very good at keeping all of that looking consistent, and completely uninterested in whether it says the same thing.
So the date gets updated on one asset and not the others, a price changes in the deck but not in the ad, the call to action drifts a little each time somebody rewrites it, and by launch a single campaign is quietly making four different promises. The person who notices is almost always a customer, which is the worst possible QA process.
An agent that reads every asset in a campaign and flags where the meaning drifted, rather than where the pixels differ. Detection turns out to be the easy half. The hard part is knowing what is even supposed to match, because text length is meant to vary between a billboard and an email, and a price is not.

Every element gets classified into one of three tiers, and the tier decides how strict the agent is allowed to be. That classification is the whole product; everything else is plumbing around it.
Once something is classified, one-click Sync All propagates the canonical value everywhere it is wrong, which is the moment the tool stops being a report and starts being useful.



Semantic role classification decides which tier an element belongs to in the first place. Sentence embeddings then handle meaning-level comparison instead of string matching, which is what stops Ends Friday and Last day Friday from being reported as a conflict, and perceptual hashing covers the hero imagery.
The Must-match tier deliberately does not use a model at all. Whether two prices are the same is not a question that benefits from judgment, so that tier runs on rules, and knowing where to stop using the clever thing is most of the design.
React · TypeScript · Hono · Neon Postgres · Claude API


Alongside the prototype: a 24 week ship plan broken into phases with what lands in each, an A/B test design with power analysis covering what to measure and how many campaigns before the result means anything, and competitive positioning against Frontify, Bynder and Canva own Brand Assist, including the places where this genuinely overlaps them and the places where it does not.

This is a working prototype and a spec, not a shipped feature. There is no usage data and no adoption numbers, and I am not going to invent any, because the entire point of the exercise falls apart the moment I start making numbers up.
What I would want next is the one thing I could not do from the outside: put the drift detector in front of ten real campaign teams and find out whether the Should-match tier is genuinely useful or just noisy enough that people start ignoring every flag. My honest guess is that the taxonomy survives and every threshold in it is wrong.
Case study · Growth prototype
Nobody asked me for any of this. Lora is an AlleyCorp-backed consumer identity app and I built the growth prototype, tested it on strangers for two months, and sent the founder my read on their product strategy. Unsolicited, all of it.
A reading you get from a name and a phone number. That is the entire input. No signup, no download, no onboarding, sixty seconds start to finish, and nothing stored anywhere afterwards.
The part that took the time was not the interface, which is frankly rough. It was the engine underneath, built on my own understanding of cycles and timing rather than wired up to a generic numerology API, because the generic version gives you generic output and nobody has ever gone quiet reading generic output. It surfaces things about a person that would normally take a full chart consultation to reach, and it was hard to build precisely because there was nothing to copy from.
Python · Vercel


Over about two months I put it in front of close to a hundred people. Friends first, then people I met at events, then strangers in coffee shops around the city, which is a genuinely uncomfortable way to spend a Saturday and the only way to find out what I needed to know.
What surprised me was not that people liked it. It was how consistent the reaction was. They would read the result, go quiet for a second, and then ask how did it know that. Same three beats, over and over, from a name and a phone number.
Months later I ran it again with a new batch of people and the reaction had not changed at all, which is the closest thing to a signal you get at that sample size.
Lora had the problem a lot of consumer apps would happily trade for: real depth and no width. Six thousand Discord members inside three months and newsletter open rates above fifty percent are the shape of a product the people who found it genuinely care about. What it did not have was a surface that brought new people in.
The obvious answer is a frictionless loop, and the obvious answer is half wrong. Removing friction does not make anybody share anything. People share the thing that made them go quiet. So the loop is worth building only if the reading at the end of it is good enough to produce that pause, which makes reading quality a growth investment rather than a content one, and that reframing is the whole argument.
Asking a stranger for their full name and phone number on a cold link is an enormous ask, and no amount of interface polish fixes it. I built the thing to store nothing at all, which helps, except you cannot prove that to someone in the two seconds before they decide.
I tried to solve it by turning up at their office in person, on the theory that trust transfers through a face and does not transfer through a landing page. That is not a scalable growth channel and I knew it at the time. It was the honest version of the problem though, and the sort of thing that only shows up when you make a hundred real people use your thing in front of you.
Along the way I formed a guess and sent it to her, because she was the only person who would actually know. My read was that Lora had deliberately kept AI out of the reading pipeline. The ritual only works if a real astrologer is doing the reading, and dropping a model into the middle of that breaks the exact trust the whole app runs on.
If that was the call, it is the harder version of the business to build and almost certainly the right one. I said as much. Knowing where NOT to put the model is a product decision most teams get wrong in the other direction, and I wanted to understand how a founder reasons about that line, which felt like a more useful thing to ask about than a role.
Acquisition without retention is a leaky bucket with better marketing, so the pitch carried four ideas aimed at depth rather than reach:
Underneath all four is the same claim: the moat is not the data, because data is copyable. The moat is that people trust the routing, and trust compounds. More precision means more trust, more trust means more engagement, and more engagement makes the matching better. That argues for spending on precision rather than reach, which is the opposite of what most growth work does.
A working prototype, a hundred conversations, interactive decks personalised for two people on the team, and no reply. The link is still live.
The number I never got is share rate broken down by reading quality, which is the one that would prove or kill the whole argument, and getting it needs volume I did not have. Everything above is a strong hypothesis with a hundred data points behind it, and I would rather say that than dress it up as a result.
Case study · Prototype
Also unprompted. Built on real Kumon data, 55 locations across the NYC metro, as a product exercise for Eulerity marketing automation platform.
Franchise marketing software is sold to corporate and used by franchisees, and those are two entirely different people with two entirely different incentives. Corporate is buying coverage across a network. The person who owns one location wants to know whether the money they spent last month did anything at all.
When the tool cannot answer that second question in about five seconds, the owner stops opening it, and a platform nobody opens generates no data, which makes the renewal conversation harder, which is how genuinely good software dies quietly.

Next.js · Vercel


Then the strategy document, four bets ordered by how much adoption each one unlocks rather than by how impressive each one sounds:

Positioned against SOCi, Scorpion and Yext, mapping where each of them already wins and where the specific gap sits. The short version is that they all compete on breadth of channels, and not one of them competes on whether the franchisee logs in.
A working dashboard on real data and a strategy document, not a shipped product. If I could only test one of the four bets it would be the Scorecard, because it is by far the cheapest to build and the entire adoption argument rests on it. If a single number does not move weekly logins, the other three do not matter.

Build note · Desktop and web app
A teleprompter, sticky notes, Claude, and a full browser, all floating on top of whatever you are doing and invisible to everyone on the call but you.
A teleprompter that scrolls horizontally or vertically, mirrors for camera rigs, and has a screen share background mode so your audience sees a clean surface while you read. Sticky notes you can drag, resize and colour. Claude, running inside a note, so you can ask for a rephrase mid-sentence. And a real browser inside another note, using an Electron webview rather than an iframe, so Google and docs actually load instead of refusing to.
Clicks have to pass straight through. Once the overlay is genuinely click-through and the hotkeys stay live while another app has focus, the thing stops being a window you manage and becomes something closer to a heads up display.
That is the whole difference. Every other teleprompter is a window, and the moment you click it to hit play you have left the app you were demoing, and the recording shows it. Getting to true transparency is why there is a desktop build at all.
The browser version runs anywhere with no install and does the teleprompter and the notes. Click-through, always on top, global hotkeys, the Claude notes and the embedded browser need real window access, so those are desktop only.
The comparison table on the site says exactly which is which, because the fastest way to lose someone is to let them download a thing expecting a feature it cannot have.
Electron · React · Claude API, bring your own key
Build note · Agent workspace
A workspace where six agents each own one job a PM would otherwise do by hand, and stream their thinking while they do it.
Competitive intel, which plans its own queries and comes back with a threat assessment. Market research, which sizes a space and maps who is in it. Feedback triage, which turns a pile of raw comments into themes with counts and sentiment. Then a PRD writer, a sprint planner and a design brief writer that consume what the first three found.
They are separate on purpose. One agent asked to do all six does all six badly, and worse, you cannot tell which step went wrong when it does. Splitting them is what makes the output debuggable, and it also mirrors how the work actually divides.
Around them sit the things that make it a workspace rather than a demo: a pipeline view, a roadmap, a repo explorer and a decision log.


Each agent emits its phase as it goes, planning, searching, writing, over server sent events. That was not a UI flourish. A research agent takes long enough that a spinner makes you assume it has hung, and the phases are the only honest way to show that it is doing something rather than nothing.
It also makes the thing debuggable. When a brief comes back thin, you can see whether it planned bad queries or found good sources and wrote them up badly, and those are completely different fixes.
FastAPI · Next.js · Tavily · Groq

Build note · Shipped product
Signal overlays and scenario testing over market data, for people whose job is to form a view and defend it.
A terminal where you put alternative signals on top of price and test what would have happened. The audience is market researchers, analysts and investors, which matters because that audience does not want to be told what to think. They want the inputs exposed so they can disagree with them.
Everything the platform draws is derived from data points a user can inspect. If an overlay cannot be traced back to something concrete, it does not ship.
Not the charts. The honesty of the outputs. It is trivially easy to build a research tool that flatters whoever is using it, because a backtest will tell you almost anything you want to hear if you are careless about how the test is constructed.
Most of the hard work is on the boring side of that line: what a number is allowed to claim, and what has to be labelled as untested. It is the least visible work in the product and the only reason anyone should trust it.
Next.js · FastAPI · Postgres · Railway
Build note · Research model
A quantitative turning point outlook for any time series. It finds the swings that matter and projects harmonic levels forward from them.
It detects the significant swings in a series and projects harmonic levels forward from them, plus a read on where the current move sits inside a larger one. The maths underneath is proprietary and stays that way.

Each of these was projected from structure that already existed before the date it points at, then checked against what the market actually did.
Those are individual calls, so here is the run behind them. On the weekly, across thirteen instruments, it caught 88% of the major turns, from 78% on JPM up to 100% on both the DOW and the S&P, with median error on tops between 0.1% and 0.8%. On the monthly it holds: MSFT and Visa at 100%, the DOW and Amazon at 86%, Apple at 83%.
Weekly is where the model lives. Daily is noise on single names, and knowing which timeframe a method is actually good at is most of the work.
The edge is real and the research is still running. Both of those are true at the same time, which is why it stays private rather than shipped.
This is the distinction I care most about in my own work. Three Axis is a product because the things it claims are things I can stand behind. Cycle X is a research model because it is still being tested, and the honest move is to keep those two words apart rather than let one borrow credibility from the other.
Build note · Consumer app
You set a daily thing you want to do, you challenge people you know, and the social cost of dropping out does the work that willpower usually fails at.
Every habit app I have used treats the problem as one of tracking, so it hands you a streak counter and hopes the number makes you care. It does for about eleven days. A streak is only accountable to you, and you will always forgive yourself.
The bet here was that a friend who can see you drop out is a stronger mechanism than any chart. Shipped the landing page and the app shell.
TypeScript · Vercel

Build note · Consumer app
A studio that builds a focus soundscape for you instead of handing you the same playlist it hands everyone else.
Focus audio is a solved problem in the sense that there are ten thousand playlists and an unsolved one in the sense that none of them are yours. The interesting version is generating the tones, because then the thing can be tuned to the person and the session rather than picked off a shelf.
Built the studio and the site. It works, and I use it, which is the only user number I am willing to claim for it.
TypeScript · Web Audio · Vercel
Build note · Retrieval agent
A RAG agent over an S3 backed document corpus, built on AWS Bedrock, with citations surfaced inline in the answer.
Operational teams were spending about twenty five minutes per request hunting through documents for an answer that existed the whole time. That is not a knowledge problem, it is a retrieval problem, and retrieval is the thing this class of system is genuinely good at. It came down to around five.
Inline citations. An answer without a source is worse than no answer in an operational setting, because a confident wrong one gets acted on and nobody can trace how it happened.
Surfacing the source next to each claim changes what the tool is. It stops being an oracle you either trust or do not, and becomes a fast way to get to the paragraph you were going to read anyway. That framing is also what made people comfortable enough to use it.
Life
Aug 10th - > lets push till 30th with full force :) Planning to switch to a new role!!
ai · general
I had spent about a week on the homepage, fussing over the copy, moving one section above another and then moving it back, doing the thing where you reread your own paragraph nine times and each time it means slightly less. Then I fetched the page the way an AI crawler would, stripped the scripts, and counted the words that were left.
Six.
Not six paragraphs, six words, and most of them were in the title tag. Here is the part nobody puts on the box: the big AI crawlers do not run JavaScript. GPTBot, ClaudeBot, PerplexityBot, they take whatever HTML the server hands them and read that. They are not sitting politely in the corner waiting for React to wake up and assemble the page, which means a site can rank perfectly well on Google, whose renderer does run JavaScript, and be functionally invisible to the thing your buyer is actually asking.
I wrote a little counter to track progress, added a noscript block with the real pitch in it, ran the counter again, and it told me 364 words. Genuine relief. I remember thinking that was a cheap fix and feeling quite pleased about the afternoon.
Except real content extractors throw noscript straight in the bin, so the number that made me feel better was measuring the file rather than what anybody actually sees. I had built a tool whose only real function was to measure my own optimism, and then I had trusted it for about four days.
Prerendering fixed it properly and took the homepage from 6 crawler-visible words to 1,860, which is the number I would put in a deck. But the lesson that stuck is smaller and more annoying than that: if you built the measurement and you built the thing being measured, check the measurement first.
What makes this whole category of bug so miserable is that the failure mode is a valid, well formed, completely empty document. Nothing throws. Builds pass, deploys go green, and the page looks perfect in your browser because your browser runs the JavaScript that the crawler never will.
So I wrote a build guard that fails the build if any page drops below a word floor, felt extremely clever about it for roughly a day, and then watched it silently kill four production deploys in a row on a different site. One codebase was building two products and I had held both to the same floor. It took eleven hours for anyone to notice, because the project being watched was the one that was fine.
Fetch your own page with a crawler user agent, strip the scripts, count the words, then open the same page in a browser and count again. The gap between those two numbers is your invisibility and the whole exercise takes two minutes, which is roughly two minutes more than most people have spent on it.
And treat a status code as evidence rather than a diagnosis. A single page app returns 200 for every URL you can think of, so anything that checks whether a file exists by looking at the status code will happily confirm the existence of files you have never written, which is how I ended up with an audit crediting me with a robots.txt that did not exist.
healthcare · general
The feature was simple enough to describe in one sentence. Someone calls a clinic, the agent asks for their insurance, and before the call ends it tells them their plan is active and their copay is twenty five dollars. That is it. That is the whole thing.
It took two days and one genuinely embarrassing reversal to work out what to buy, and both detours were the same mistake wearing different clothes.
I found a YC company that reverse engineers login walled portals into stable APIs, which is properly smart work, and their pitch line is one of the better ones I have read: UIs are for humans, not agents. I got excited on the train home and started sketching the integration in my head before I had checked anything at all.
Then I said the requirement out loud to myself, which is a habit I should use more often. There are hundreds of payers. Their model wraps one portal at a time. I was about to hand build, vendor by vendor, the exact problem that clearinghouses have existed to solve since before I could drive, and I was going to feel innovative doing it.
Their site also had no mention of HIPAA or a business associate agreement anywhere on it, which for anything touching patient data should have been the first question I asked and was somehow not in my first five.
Having dodged that, I then ruled out the obvious modern option because I remembered it having a five hundred dollar monthly minimum, which is a lot of money to commit to a feature that might not work. Except it did not have one, and had not for a while. I was making an architecture decision from a memory of a pricing page rather than from the pricing page.
The real number, once I bothered to open the tab, was thirty cents a check falling to eight at volume, no minimum, free sandbox. At the volume I actually had that came to about twelve dollars a month. I had nearly designed around a constraint that no longer existed, which is a more expensive habit than being wrong about something you never knew.
Somewhere in the middle of all that I found the sentence that explains the entire market. Where an electronic rail exists, calling is dead. Where it does not, calling is the product.
Eligibility has a rail, because a law forced one into existence, so it costs cents and comes back in seconds and nobody has built a company around phoning payers to ask whether a patient is covered. Prior authorization has no rail, so it is portals and faxes and forty minutes of hold music, and there are venture backed companies whose entire business is putting an AI on hold so that a human does not have to be.
Same industry, opposite answers, and the only thing separating them is whether somebody, decades ago, sat down and wrote a standard.
ai
Everyone who builds one of these has the same first day. You wire speech to text into a model into text to speech, you call your own phone, it picks up and talks back, and for about ten minutes you feel like a wizard. I sat in my kitchen having a full conversation with something I had assembled that morning and genuinely could not stop grinning.
Everything between that afternoon and a clinic putting it on their main line is the actual job, and almost none of it is the part that felt like magic.
Endpointing is deciding when the caller has finished speaking as opposed to pausing to think, and it sounds like a solved problem right up until you watch it fail on somebody real. People pause in the middle of exactly the sentences you most need to get right, and no amount of clever prompting fixes a system that has already started talking.
An older patient reading a member ID does not say W123456789 in one breath. They say W, one, two, three, and then they take a breath and check the card, and then four, five. Set the silence threshold tight and your agent interrupts a 78 year old halfway through a digit. Set it loose and every conversation carries that satellite delay that makes people say hello? twice.
Voice activity detection handles the easy half of this. The rest needs a model making a judgment call about whether a sentence is finished, which is a genuinely strange thing to be tuning at eleven at night, and which I now think is the single hardest problem in the whole stack.
Word error rate on clean read speech is the metric everybody quotes and it is close to meaningless, because nobody calls a doctor in clean read speech. They call from a car, from a waiting room, holding a toddler, halfway through a sentence they started before you picked up.
The slice that decides whether a booking is correct or quietly wrong is narrow and unglamorous: spelled out alphanumerics, dates of birth, payer names, and callers switching between English and Spanish inside a single sentence. Measure there and your headline number gets noticeably worse, and for the first time it starts telling you something true.
Anyone can demo one of these. The people who have run one in production for a while all end up talking about the same thing, which is replay suites: real recorded calls, re run against every prompt change, before it ships. A prompt is code, and you would not push code without tests just because it read nicely on the screen.
The measure that matters is task success, meaning the appointment actually landed in the calendar, not whether the transcript reads well, because transcripts read well all the time while nothing happens. And tool calls get tracked separately, because an agent that says every right word and then calls the wrong function is far more dangerous than one that fails loudly.
The loud failure you fix on Tuesday. The quiet one you hear about from a patient.
general
The report opened about as strongly as a report can. Google Analytics running on a healthcare site, a HIPAA violation, fifty thousand dollars in potential fines, all of it in a red box at the top of page one. I read it twice on my phone and felt that specific stomach drop where you already know the rest of your evening is gone.
Then I grepped the served HTML, and there was no analytics on the site. Not misconfigured analytics, none at all, not a single tag. The most urgent finding in a paid audit was invented.
The same report credited me with six files and three legal pages that also did not exist, and once I worked out why, it stopped being infuriating and became sort of funny. A single page app serves index.html for any URL it does not recognise, so every path on the site returns HTTP 200, including paths for files nobody has ever created.
The scanner asked whether robots.txt existed, got a 200 back, ticked the box, and reported a file I had never written. It was not lying exactly. It was measuring a shell and describing a house.
Now I ask how a number was produced before I believe it, and in that same week the habit reversed three separate conclusions. Two API keys came back 403 and the obvious read was a permissions problem, until I looked at the response body and found error code 1010, which is a bot block on the user agent that fires before authentication is even attempted. With a browser user agent both keys worked perfectly.
A search tool showed my site for no query at all, so I announced it was not indexed, and it turned out to be ranked first for the query I cared about. The tool was the problem. And a set of new pages looked completely fine while rendering their headings in the body typeface, because a wrapper class was missing, which taught me that looks fine and verified are different words.
While I was in there I found the real problem, and it was not in any of the pages the audit had looked at. Live on the marketing page, for a company with no customers, sat a compliance badge for an audit that had never happened, traction numbers for users who did not exist, and testimonials from named clinicians who had never heard of me. Placeholder copy that had quietly become production copy while I was busy elsewhere.
That is worse than any score, and not for the reason you would guess. Someone who signs a contract relying on a certification you do not have has a fraud claim, and fraud voids the liability cap in your own terms. The fake badge was not just embarrassing, it was CANCELLING THE LEGAL PROTECTION written on a different page of the same website.
I deleted all of it that night.
ai · healthcare · finance
A medical biller calling an insurance company is on the phone for thirty or forty minutes, and roughly six of those minutes involve talking to another human being. The rest is a phone tree, a hold queue, and a saxophone that has been looping since 2011.
For a person that is ten or twelve dollars of labour, but the real cost is that it is exclusive attention. One human, one call, no multitasking, and hold music priced at exactly the same rate as conversation.
For an AI agent the cost splits in two, and only one half of it is expensive. The phone line has to stay open the entire time, which runs about a penny a minute, but the models only do work when somebody is actually speaking, so you pay for six minutes of inference rather than forty minutes of waiting.
That puts a forty minute hold at roughly a dollar of machine time against twelve dollars of human time, and the ratio gets better the longer the payer keeps you waiting. Which is the genuinely interesting part: hold time was the lever insurers used to ration access to their staff, and it quietly stopped working the moment the caller stopped being bored.
Concurrency is doing even more work than the cost per minute. A human biller manages maybe ten or twelve payer calls in a day because holding consumes the day, while one system sits on five hundred holds simultaneously without getting bored, without batching them for Friday afternoon, and without quitting after six months of saxophone.
That last one matters more than it sounds, because turnover in those roles is brutal and entirely understandable.
This is a melting iceberg business. Every payer interaction that gets a real API permanently deletes some call volume, and the regulatory push is heading that direction, so the addressable pile of hold music shrinks a little every year.
It is a very large iceberg, mind you. Healthcare still runs on fax machines in 2026. But anyone selling you a ten year thesis built on hold music is selling you something.
quant · finance
I have built two things in markets, one of which is a research model and one of which is a product, and for a while I treated them as the same object at different stages of completeness. They are not. Confusing them is the most expensive mistake available in this corner of the world, and almost everybody makes it about ten minutes after a good backtest.
A backtest is a hypothesis with better graphics. It tells you that a rule would have worked on data you already have, which is a far weaker claim than the equity curve makes it feel at one in the morning when the line goes up and to the right and you start mentally spending the money.
The failure that actually scares me is look ahead bias, where information leaks backwards into a decision the model could not have made at the time. It is easy to introduce, nearly invisible afterwards, and the giveaway is that the output looks fantastic. A result that looks great and cannot fail is usually not a result.
The only honest test I know is to run the thing on noise. If it prints a beautiful curve on randomly generated data then you have not found a signal, you have found a bug with very good manners, and the difference matters a great deal more to your users than it does to your ego.
A research model owes you an honest confidence interval and nothing else. A product owes you something different, which is repeatability, and the discipline not to overclaim while you are still working the rest out. Those are almost opposite instincts, and holding both at once is most of the job.
Which is why the platform I actually ship gets described as signal overlays and scenario testing and carries no accuracy claim anywhere on it. Not because there is nothing there, but because the number I could put on a marketing page is precisely the number I would least be able to defend in a room with someone who does this for a living. Adoption I can defend, load times I can defend, and prediction accuracy has to survive a stranger with a grudge and a Jupyter notebook.
Write down what you expect before you run the test, and write down what would make you wrong, because the moment a result is thrilling you lose the ability to tell those two things apart. Do it in advance and the thrilling result arrives already labelled.
I have found real things and I have tuned until the graph got pretty, and I promise you they feel identical from the inside. That is the whole problem.
life
I ship software all week, which means I spend most of my waking hours inside systems where every problem is theoretically solvable by thinking harder at a screen, and where thinking harder at the screen is therefore always the available move. Running is the one thing I do where that move does not exist.
Weekday problems are legible. A bug has a cause, a metric has a definition, a payer has a phone number, and if you sit with any of them long enough they eventually give. The other kind of problem, the one about whether you are building the right thing at all, does not respond to sitting with it. It shows up about forty minutes into a run, once you have finally stopped narrating your own week to yourself.
Most of my real product decisions have arrived somewhere around mile four, and never the implementation ones. Always the layer above: the quiet oh, we are solving the wrong problem, which has not once turned up while I was staring directly at the repository containing the wrong problem.
I host small events around the neighbourhood and I try every cafe that opens, which as a vegetarian in New York is a sport with a high failure rate and an unreasonable amount of cauliflower. Both of those are really the same habit as the running: get out of the room, talk to somebody whose job you could not do, and come back carrying a question you did not leave with.
Nearly everything I have built started as someone describing an annoyance I would never have run into on my own. It is not a networking strategy and I would be embarrassed to call it one. It is just the least efficient and most reliable idea generator I have found.