Skip to main content

Featured

My AI Wants to Go Shopping. Retailers Would Like to Know Whose Side It’s On.

I used to think the most complicated part of shopping was figuring out whether I actually needed something. Apparently, that was an embarrassingly primitive concern. Now I am supposed to consider whether an artificial intelligence agent should evaluate my needs, compare the options, negotiate the digital obstacle course, and spend my money while retailers argue over whether to let it through the door. We have finally reached the stage of civilization where buying socks requires a debate about sovereignty. The headline that set me off was The Wall Street Journal’s “Retailers From QVC to Kohl’s Are Split on Agentic Shopping.” According to its October 2 reporting, QVC and HSN welcome purchasing agents, Kohl’s prevents agentic purchases, and Tapestry allows browsing but not buying. Retailers’ concerns include customer experience and control over data. Those are the reported positions; what interests me is the argument underneath them. Everyone wants to help me shop, but I suspect we have d...

Claude Has Entered the Lab, and Apparently I Still Have to Think


I read the phrase “Claude-shaped science” and immediately pictured a laboratory being remodeled around a chatbot’s preferences. The benches would be rearranged, the microscopes would receive motivational feedback, and someone with an expensive haircut would announce that curiosity had finally achieved product-market fit. I realize that is an unfair place to begin. Unfortunately, I have spent enough time around technology announcements to develop an involuntary twitch whenever a company’s product name starts attaching itself to a fundamental human activity.

Then I read the actual piece, and something inconvenient happened: I found the central idea interesting. In his October 1 guest essay for Anthropic, physicist Matthew Schwartz describes building BootLoops, an open-source scientific toolkit, around problems that suit current AI capabilities. His account includes an important complication: technical achievements sometimes needed specialists to redirect them toward questions worth answering. He also discloses that he was a visiting researcher at Anthropic during the project. That context matters, although it does not settle whether the work is valuable.[1]

I can appreciate an argument while keeping one hand on my wallet. In fact, I would recommend that posture for most encounters with the technology industry. It allows room for real progress without requiring me to stand beneath a confetti cannon every time software completes a difficult assignment. My interest here is less in crowning a digital genius than in asking what happens to our judgment when an exceptionally capable tool becomes the most confident presence in the room.

There is something deeply appealing about the promise of assistance that arrives without the usual human complications. A machine does not need its contribution praised during lunch. It does not spend the meeting explaining why your suggestion resembles an idea it had in 2017. At least, I do not have to worry about hurting its feelings when I close the window. For anyone exhausted by the social overhead of getting things done, that alone sounds dangerously close to a vacation.

But convenience has a way of sneaking into the chair reserved for authority. I ask a system to help organize my thinking, and before long I am organizing my thinking around the kinds of responses it produces. The shift can be almost embarrassingly comfortable. Everything becomes clearer, smoother, easier to summarize. I get a tidy explanation and the pleasant sensation of forward movement. Whether I have actually understood anything becomes a separate question, which I may postpone because the formatting looks terrific.

That is the question the title leaves rattling around in my head. If we start looking for problems that fit our tools, how do we keep track of the problems that do not? I am not accusing this project of abandoning difficult research. I am asking about a temptation I recognize in myself: choosing the task that makes me feel effective, then quietly upgrading that feeling into evidence that I chose the most important task available.

Give me an afternoon with a messy obligation and I can become extraordinarily productive at something adjacent to it. I can reorganize notes, improve a title, research a better method, and construct a schedule so attractive it deserves its own little frame. Meanwhile, the actual obligation sits untouched, developing a personality. My concern about powerful AI begins there, in a very ordinary human weakness. A better machine could make avoidance look impressively professional.

Imagine a research institution discovering that one category of project now produces results much faster than another. I would expect people to notice. I would expect managers to notice even sooner. Then I can imagine the annual review forms changing, the funding language shifting, and the slower work acquiring the vaguely embarrassing status of an elderly relative who still pays by check. Nobody would need to announce that the tool was setting priorities. The incentives could handle that quietly.

I would want someone in that institution asking which questions had disappeared from the conversation. What stopped being proposed because it would be difficult to demonstrate progress? Which project needed patient observation rather than another computational sprint? I do not think those questions should function as brakes applied automatically to every innovation. I think they belong on the dashboard. A speedometer is useful, but I would also like to know where the vehicle is going.

My skepticism becomes sharper whenever output starts impersonating significance. I can count pages, charts, analyses, and completed tasks. Counting understanding is more awkward. It requires discussion, disagreement, and sometimes the humiliating discovery that a beautiful result has answered a question nobody especially needed answered. I would rather endure that embarrassment early. Otherwise, I risk becoming the proud owner of a magnificently engineered staircase that reaches the wrong window.

This is where I find the human side of the story compelling. I want expertise to remain capable of being unimpressed. That sounds like a small request until I picture the social pressure surrounding an expensive new system. Someone has approved the budget. Someone has promised transformation. Someone has already prepared a presentation using a photograph of a glowing brain. Into this atmosphere walks a person whose contribution is to say that the result needs more work. I want that person protected.

I do not mean protected from criticism. I mean protected from the assumption that caution signals incompetence or hostility. When I hear hesitation treated as a character flaw, I start wondering what the organization actually rewards. There is a difference between enjoying a promising result and demanding that everyone enjoy it on schedule. I can live with a colleague who asks irritating questions. I am much less comfortable with a culture that makes those questions professionally expensive.

At the same time, I have no interest in defending drudgery as a sacred educational experience. If a tool can eliminate hours of mechanical work, I am happy to let those hours go. I do not need researchers to suffer through repetitive tasks so that the eventual insight feels properly earned. My toaster does not diminish breakfast by sparing me an intimate relationship with fire. Difficulty deserves respect when it teaches something, rather than merely consuming somebody’s Tuesday.

The hard part, for me, is telling useful struggle from disposable friction. Learning to notice a suspicious result may require wrestling with calculations. Learning to move information between incompatible formats may mostly require patience and a deteriorating attitude. I would like educators to think carefully about that distinction. Saving time is wonderful, but I would hate to discover that we removed the experiences through which people learned to recognize when something had gone wrong.

I imagine a young researcher being handed a polished analysis that would once have taken weeks to assemble. The obvious question is whether the analysis works. My next question is whether the researcher can explain why it works, where it is fragile, and what would make them distrust it. I would ask those questions without turning the exercise into a purity test. Using assistance should be normal. Being unable to interrogate the assistance should make us pause.

I feel a version of that tension whenever a sentence arrives more elegantly than the thought behind it. Smooth language can make me believe a loose idea has finished developing. I read it back, admire the rhythm, and almost forget to check the claim. That vulnerability is mine; blaming the software would be wonderfully convenient. Still, I want tools and working habits that expose uncertainty instead of applying a fresh coat of verbal paint over the cracks.

If I were reviewing an AI-assisted project, I would want to see some of the mess. Show me the assumptions that changed, the promising approach that failed, and the decision that required actual debate. I would learn more from that record than from a triumphant demonstration conducted under lighting normally reserved for luxury automobiles. A credible account should help me understand how the people involved caught mistakes. Perfection presented without a trail makes me curious for all the wrong reasons.

I would also want to know what the whole effort cost. My question would include money, time, supervision, and the attention required to verify the result. A calculation that takes minutes can still sit inside a workflow that consumes days of human checking. That might be an excellent trade. I simply want the trade described honestly. I have reached the stage of life where the phrase “it only takes five minutes” causes me to check whether somebody has excluded the preceding week.

The same accounting should include work that went nowhere. Suppose I hear about several dazzling successes from a much larger set of attempts. I want the larger set in view. Otherwise, I cannot tell whether I am seeing a reliable method or the winners selected for the family Christmas card. Failure would not automatically undermine my confidence. A clear picture of failure could increase it, because then I would know which expectations the evidence actually supports.

I am especially wary of demonstrations that let the audience supply the most extravagant conclusion. Nobody explicitly promises that the machine understands everything. They simply show an impressive result, pause meaningfully, and allow our imaginations to sprint toward the Nobel ceremony. I recognize the technique because I am susceptible to it. I like stories with decisive turning points. A long process full of conditional improvements asks more patience of me than a sudden arrival from the future.

Yet my favorite possibility here is relatively unglamorous: more people having enough time to examine things properly. I would love a world where researchers could check additional explanations because preliminary work became cheaper. I would love a student to pursue a modest question without first assembling an empire of resources. Those are hopes, not outcomes I can certify from one article. They also strike me as more humane ambitions than asking how quickly we can remove the humans from the photograph.

That distinction matters because I am tired of discussing capability as though employment were a game of musical chairs staged for spectators. Whenever a tool improves, the conversation seems eager to locate the next person who should become unnecessary. I would rather ask what valuable work that person could now do. There will be real disagreements about resources and roles, but I refuse to accept that the most imaginative use of intelligence is always a shorter payroll.

For me, the promise of assistance should include better working conditions. If a research group can accomplish a task faster, perhaps its members could spend more time thinking, collaborating, or sleeping like organisms that require maintenance. I can already hear the objection that competitive pressure will absorb the savings. That is precisely why I want the question asked early. Otherwise, we may build astonishing tools and use them to make everybody late for dinner more efficiently.

Access raises another set of questions I cannot wave away with enthusiasm. If sophisticated AI becomes central to certain research workflows, I want to know who can afford reliable access and who has to ration it. I also want to know what happens when prices, terms, or product behavior change. A laboratory planning years of work needs more than the reassurance that a service currently has a friendly interface. Dependence deserves an actual plan.

I would prefer institutions to preserve room for alternatives, document what their workflows require, and avoid letting familiarity become a permanent purchasing decision. That preference is not peculiar to AI. I feel the same way about any tool that becomes difficult to leave. The honeymoon phase of software adoption is full of possibility. The interesting questions arrive later, when the invoice grows teeth and everyone discovers that exporting their work requires an afternoon and a minor spiritual crisis.

Branding makes that dependence easier to overlook. A memorable name gives a complicated system a personality I can address. Before long, I may talk about it as if it were a colleague with stable intentions and a recognizable temperament. I understand the appeal. Names are useful. But I want to remember that a product relationship can change through a release, a policy decision, or a pricing update, without the product having a difficult conversation with me beforehand.

There is also a question of credit that I would want handled generously and precisely. When a result emerges from datasets, software, expert interpretation, and many rounds of revision, I want the account to acknowledge those contributions. The most visible interface should not swallow the whole story. I am impressed by tools that make knowledge easier to use. I am also impressed by the people who spent years producing knowledge that somebody else can now access in seconds.

I think about the hypothetical technician whose careful records make an analysis possible, or the researcher who maintains an unfashionable dataset long after the original excitement fades. In my preferred version of the future, those people become easier to recognize. In the version I worry about, their labor becomes scenery behind a dramatic claim about automation. I would like us to resist telling stories so streamlined that the people who made the result possible disappear from them.

My impatience with hype should not be mistaken for a desire to see the technology fail. I want useful discoveries. I want difficult questions to become tractable. I would be delighted to have my cautious expectations overtaken by convincing evidence. What I resist is the suggestion that enthusiasm earns a discount on scrutiny. If a capability is as consequential as its advocates believe, I would expect them to welcome serious examination of where it works and where it breaks.

I also owe the same fairness to the skeptical side. Dismissing every success as mere automation would let me preserve my worldview without doing much thinking. That is an attractive bargain and a terrible intellectual habit. A tool can be limited and still change what people can accomplish. I want to be specific enough to notice that change. Rolling my eyes is easy; explaining exactly which claim I doubt requires me to put the eye muscles on leave and use something else.

The standard I keep returning to is accountability. If I rely on an AI-generated analysis, I remain responsible for the decision to rely on it. I do not get to point toward the chat window as though an unusually articulate weather event passed through my office. That responsibility should shape how I work before anything goes wrong. I want records, checks, and collaborators who can challenge my interpretation while there is still time to correct it.

I would treat the system’s confidence as an invitation to investigate, rather than a substitute for investigation. I would ask what evidence supports the result and what evidence could overturn it. I would want independent checks when the stakes justify them. None of this requires me to adopt a posture of permanent suspicion. I check a map because I intend to travel. Verification is part of making a tool useful enough to trust with a real destination.

Perhaps the most uncomfortable part is that this asks more of my judgment precisely when the technology promises to make life easier. I may spend less time producing a first answer and more time deciding which answer deserves attention. That is work too, and I should resist pretending otherwise. A screen full of plausible possibilities can be liberating, but it can also resemble a restaurant menu with eight hundred entrees and no indication that the kitchen owns a stove.

I would extend that concern to the people reading the eventual coverage, including me. A headline can send an impression halfway around the world before the qualifications have found their shoes. When I write about an impressive system, I want to make the boundaries visible in the main story. Otherwise, I am helping manufacture the very confusion I claim to dislike. There is no special virtue in burying the careful sentence near the bottom and hoping readers bring archaeological equipment to their morning browsing.

I also want room for delight. Sometimes a clever connection or an unexpectedly effective method deserves a moment of uncomplicated appreciation. I do not want to become the person standing beside every achievement with a clipboard and a damp towel. The challenge is letting amazement remain an emotion instead of promoting it to a finding. I can be excited by what might follow while admitting that I do not know what will follow. That admission leaves the future open without making my expectations somebody else’s evidence.

And if the work eventually changes an entire field, I hope the story includes the awkward middle. I want to hear about the disagreement that improved the question and the check that spoiled a beautiful answer. Those details would help me understand how to use powerful assistance responsibly. They would also make the people involved recognizable. I have little use for a future populated exclusively by visionaries announcing inevitabilities. I would much rather meet a few competent adults who can explain what surprised them and what they still cannot figure out.

I find myself wanting a quieter vocabulary for progress. Tell me which task became possible, how the result was checked, and what remains uncertain. Let the accomplishment occupy its actual size. If it is substantial, it will survive the absence of a civilization-ending or civilization-saving headline. I do not need every useful development to arrive wearing a cape. Some of the improvements I value most would look disappointingly ordinary in a promotional video.

Reading about Claude-shaped science leaves me hopeful, but my hope has paperwork attached. I want tools that expand the questions people can pursue, institutions that reward careful judgment, and enough intellectual independence to disagree with a polished answer. I would happily welcome Claude into the lab under those conditions. I would even give it the repetitive work. I would keep asking who chose the question, who checked the result, and whether we learned anything worth knowing. The machine can have a seat at the table. I am keeping my obligation to think, even when the interface is charming.

Source: [1] Matthew Schwartz, “Claude-shaped science,” Anthropic, October 1, 2026. This post is independent commentary on the essay, not independent verification of its reported research results.

Comments