SCROLL NEWS / DISCOVERY
Search headlines
Find the stories shaping the conversation.
Results for " Claude Code " 45 found 🔀 AI-powered Shuffle
AI Makes "Open-Source, Clean Room" Implementations Of Adobe Photoshop & Premier In Rust
As an interesting demonstrator for the power of AI and potentially pressing legal challenges, Claude Code and AI agents have created 'open-source, clean-room reimplementation' versions of Adobe Photoshop, Adobe Premier, and Adobe Lightroom
Claude Code relaunches Projects to manage multiple AI agents in the cloud
Claude Code can run teams of AI agents in one spot now.
Anthropic launches Claude Code Projects, an ‘always-on’ conversation that remembers and delegates your long-running dev work
For enterprises, it makes a whole lot of sense: their digital storefront, website, content management system, procurement platform, or other business application rarely has a discrete endpoint.
Anthropic's Labs team drives AI innovation as IPO nears
Anthropic Labs built some of the company's biggest hits, like Claude Code. Cofounder Ben Mann lays out his team's efforts and changing ambitions.
AI needs science’s search history
“Actually, that’s where the gold is,” said Alasdair Russell, my graduate school friend who leads a pre-clinical genome editing group at Cambridge. He was talking about the winding road of science that is omitted from published papers. “When you’re discussing how to do the experiment, and why this way is better than another way, and what does the data really mean? I know what it shows, but what does it mean?” At RAAIS 2026, he described how his group has begun logging the twisting path of discovery as it happens, recording the verbal and written exchanges that normally disappear. Ideas become nodes in a living graph: they branch when a meeting produces two plausible experiments, merge when separate lines of evidence converge, and go dark when someone quietly stops pursuing them. In his implementation, one agent scores novelty, while another tries to learn “how scientists think and how they navigate through a complex world of data,” so that high-potential nodes can trigger deeper investigation. As frontier AI labs deploy agents toward scientific discovery, giving AI authentic scientific taste remains a trillion-dollar dilemma. The bottleneck is that the record we have kept for centuries of science might be insufficient to get us there. The experiments that never make the paper Peter Medawar, 1960 Nobel laureate and “father of transplantation,” asked in 1963 whether the scientific paper was a fraud, and answered yes: the form of the paper misrepresents the thinking that produced it. Sixty years on, the diagnosis is unchanged. A paper presents a clean progression from hypothesis to result to conclusion. Lost along the way are the unconventional theories, the abandoned or unaffordable methods, and the underwhelming and inconclusive data. It reads like a browser history with every dead end deleted: the ten open tabs, the four rephrased queries, the wrong turn down a subforum. All scrubbed, leaving one clean path from question to answer as the canonical path. What has changed since then is that the discarded material is now worth something. No high-impact journal wants these artifacts, but raw trial and error is the most likely source of the training data required to develop scientific intuition, or what researchers call taste. Earlier attempts to capture the discards were motivated by scientific integrity. The Journal of Negative Results in Biomedicine launched in 2002 to publish rigorous studies that disputed established models or exposed ineffective treatments. Its archive preserved negative conclusions after they had become papers, rather than the live alternatives and arguments that produced them. BioMed Central closed it in 2017, saying the mission had been served now that other journals publish null results. A less generous reading is that in fifteen years it published around 200 papers because almost nobody wants to read a negative result, let alone write one up when it will not count toward academic tenure. AI models are different. Even an uninteresting negative result can be useful, provided it is labeled. Earlier this year, Anthropic put a version of this to the test. The company pulled 129 decision points from real Claude Code sessions between January and March 2026, showed models the work up to a human detour, and asked what should happen next. Claude Mythos Preview beat the human choice 64% of the time, while Opus 4.5 managed 51%. The comparison was tilted toward the models because Anthropic deliberately chose moments where the human decision had room for improvement. On 127 further scenarios where the human action was already strong, the models improved on it only about 20% of the time. The study was possible because Anthropic’s researchers work inside a tool that logs by default. The reasoning, detours and outcome are produced in the same working environment. Biology has no equivalent. Its reasoning happens in hallways, at benches, in Slack threads and on whiteboards, while the outcome arrives weeks later somewhere else. The record has to emerge from the work itself. Ask scientists to reconstruct it afterwards and we’ll create another polished account. Subscribe A log is not a label Before any of this becomes training data, a negative result has to say what failed. The first paper published in the Journal of Negative Results in Biomedicine examined 234 negative studies across five leading medical journals. Only 30% discussed statistical power, and half clearly defined a primary outcome. Tell a model an experiment failed, without telling it whether the assay was underpowered, a reagent had degraded, or the hypothesis was simply wrong, and it will learn noise with confidence. Decision histories have a second missing label too. When a lab considers five experiments and runs one, only the chosen branch returns an outcome. The other four are experimental counterfactuals. Researchers call this the selective labels problem: the data reveal results only for the actions someone chose to take. Such a record can teach a model to imitate a lab’s taste, but it cannot establish that the taste was good. It also smuggles local constraints into general lessons. A lab that never ran cryo-EM because it did not own a microscope teaches a model that cryo-EM is rarely the right call. Alasdair’s nodes preserve the candidate set, which is the necessary first step. To be useful, each decision record needs six fields: The evidence available at that moment, and nothing that arrived later. The candidates considered, including the ones dismissed in a sentence. For each candidate: the expected result, the confidence attached to it, and the cost in time and money. The route chosen, and the reason. What happened next. How the evidence changed the scientist’s view. These must be timestamped before the outcome, because hindsight turns uncertainty into inevitability. Disagreement also has to survive. Averaging three scientists into one clean rationale reproduces exactly the information loss we are trying to fix. But there is an obvious failure mode here. Once decisions are logged and scored, people log for the record post-facto. Anyone who has watched an electronic lab notebook fill up with retrospective tidying knows how quickly a research tool becomes a compliance exercise. A record that costs a scientist ten minutes of honest reflection per decision is worth far more than one that costs an hour of performance. A publication initiative called Registered Reports offers a useful starting point. Here, researchers submit their rationale, methods and analysis plan for peer review before the data exist, and the journal commits in principle to publish if they follow the approved plan. Nature has now expanded the format across every field it covers. This approach to paper writing timestamps intent before the outcome, but it freezes one plan. A useful search history must also preserve how the plan changed, which alternatives were rejected, when, and on what evidence. Subscribe Run the runner-up Autonomous labs show what happens when decisions and outcomes are connected in a loop. Liverpool’s mobile robotic chemist ran 688 experiments over eight days in a ten-variable formulation space. A batched Bayesian search used each result to choose the next experiments and found photocatalyst mixtures six times more active than the starting formulations. Every completed experiment changed what the system did next. The robot optimized within a goal and search space that humans had already chosen. It did not decide which scientific question mattered. So what this work demonstrates is narrower, but still useful: a search history pays off when the options return standardized outcomes quickly. Open-ended biology is harder because branches can take weeks and the discarded alternatives may never be run. Some exploration therefore has to be bought. Where two branches are plausible and the stakes justify the cost, a lab should sometimes run the runner-up. Otherwise the record captures what today’s scientists usually chose and stays silent on what they systematically overlooked. A funder could ring-fence a small fraction of a grant for the branch not taken, provided that the decision record and the outcome are both deposited. That costs money, but so does having every lab rediscover the same abandoned path. Test whether taste transfers When it comes time to evaluate if our new AI system exhibits taste, the comparison should run across laboratories and fields. At selected decision points in a live research campaign, we should freeze the six fields: the evidence, the candidates, expected results, confidence, costs and the choice made. Experts then rank the options before the outcomes are known, and those outcomes are checked independently later. We could give models one of two diets - papers alone, or papers plus decision histories - and ask them to rank the next experiment on unfamiliar projects. Then, we score information gained per dollar and week, calibration, and how quickly weak branches are abandoned. If histories improve choices only inside the lab that produced them, they amount to useful organizational memory. If they improve choices on unfamiliar problems elsewhere, Alasdair’s group will have captured something science has never managed to write down: a transferable record of taste. Until that happens, a scientific search history is a promising record, not yet a training set.
Claude Code brings live iOS app testing into its Mac app
Claude Code’s Mac app now lets users test iOS apps in an interactive simulator pane, provided they have Xcode with the iOS platform installed.
Claude Code Can Now Build and Test iOS Apps in Apple's Simulator
Anthropic today said Claude Code for desktop has been updated to work with the iOS Simulator, with the integration available in public beta. Claude Code can open in-progress apps in the iOS Simulator pane when it's instructed to build, run, or check an app. Claude can watch the iOS Simulator live as it runs, interact with it, and then iterate until a project is finished. Developers can continue to use the iOS Simulator as Claude works.
This 12-year-old created an AI receptionist to help small businesses
Mana Jampala is a Gen Alpha founder who is already building enterprise tools using ChatGPT and Claude Code.
Forget prompts: 'Loop engineering' is all the rage now
Claude Code creator Boris Cherny says he doesn't "write the prompt anymore." Here's how loop engineering is changing coding.
AI optimizer beats Claude Code, Codex by 2.5x
Arbor separates strategy from execution using isolated git worktrees, so engineering teams can finally trace which optimization actually moved the needle.
Anthropic ships major Claude Design overhaul with design system imports, code round-trips, and a fix for its token-burning problem
Anthropic has overhauled Claude Design with brand-compliance controls, Claude Code integration, lower token usage and new enterprise app exports, positioning the AI tool as a serious platform for design-to-code workflows.
We can create the future of science right now
Paul Litvak is the founder and Executive Director of the Robyn Dawes Institute and a Visiting Scholar at UC Berkeley. He has a PhD in Behavioral Decision Research from Carnegie Mellon where he studied emotions and the sunk cost bias. Over 15 years in industry, he solved a wide range of challenging technical problems. At Meta he created models optimizing the ad review process. At Google he ran experiments to measure social influence. At Airbnb, he was a product manager leading machine learning teams optimizing search ranking and price suggestions. He also co-founded and led product at Rhythmic Health, a venture-backed biosensing startup, creating an accurate low cost system using a color changing strip and a smart phone to measure salivary lactate. The bottleneck At this point, it is uncontroversial to say that science needs to stop using the PDF article as the unit of knowledge and currency. The unbearable slowness of scientific publishing, the profit motive and margins: I’m not saying anything new. The PDF also sucks because it’s hard to extract structured information, which makes it hard to do evidence synthesis. As a result, we do much less evidence synthesis than is needed. And evidence synthesis ultimately undergirds most policy and medical decision-making. I can see second-by-second real-time odds for any sporting or newsworthy event, but a school board can’t see the best evidence on whether their 8th graders should be taught algebra. As a society, we don’t treat this as an important problem. Again, not controversial. Not only is the problem well understood, but the solution has already been laid out. What we need is AI-assisted living evidence synthesis - (1) an open knowledge graph of atomic claims, (2) claims linked to evidence, (3) assessment and synthesis of each piece of evidence, and (4) continuous updating with new data. What few realize (yet) is that the technical capacity to build this vision for a significant portion of science already exists. Not only that, scientists and startup teams are already building many of these components. I know this because I’ve been surveying the space and talking to many of the builders. There are some missing pieces: for example, evaluations of how well some of the components work. But at this point, most of what’s missing is a fully end-to-end working integration of all of these parts. In the rest of this essay, I’m going to lay out all the parts of a working living evidence layer for science and who is working on them, and propose concrete next steps for building this system. What’s now possible The diagram above outlines the components of a living evidence synthesis platform, including some of the teams working on each component1. Scientific PDFs are processed into claims with associated evidence. The evidence is subjected to a forensic audit, methodological evaluation and robustness and reproducibility checks. Finally it’s given a weight in a continuously updating synthesis. What follows is a description and status of each component and a few of the teams working on them. Document understanding The first thing you need to be able to do is turn an article into structured data. Mostly that means parsing PDFs. There are often multicolumn layouts that confuse non-specialized PDF to text processing libraries. For scientific papers, there is the added complexity of parsing formulas and tables and figures. This is a really hot area - there are startups offering APIs, and it seems like a new open source package gets posted to Github every few weeks. What follows isn’t exhaustive. A package called GROBID was the state of the art for a while, they didn’t update their package for nearly two years until very recently. In the meantime reducto.ai released an AI powered PDF extraction API, PaddleOCR became popular, IBM released a model called Docling, and both Mistral and Gemini created models and libraries. I also know of at least one other well-funded psychology research group working on a paper parser. By contrast, there are few open evals in this space, with no extensive evals for complex table comprehension in particular. Nonetheless, I’m confident this will be a solved problem soon, given the combination of LLM advances and developer interest. Hypothesis level extraction There has been increasing interest in comprehending the extracted text of papers and linking information to evidence for each hypothesis. A lot of work has already been done. Trialstreamer (Marshall et al. 2020) and RobotReviewer LIVE (Marshall et al. 2023) demonstrated automated extraction of trial population, intervention, and outcome at scale on clinical RCTs. PaperQA2 (Skarlinski et al. 2024) and Ai2 ScholarQA (2024) extended this to retrieval-augmented question answering with citation grounding. Elicit, Consensus, and SciSpace operationalized claim-level extraction for end users. OpenEval (Booeshaghi et al. 2026) is the most recent and most ambitious: 1.96 million atomic claims extracted from 16,087 eLife manuscripts using Claude Sonnet 4.5, grouped into ~299,000 results, with LLM evaluations showing 81% agreement with human peer review on a 2,487-paper subset. None of these solutions link claims to test-statistics, as you would need to evaluate randomized controlled trials. This is why I built the evidence.guide API - to extract hypotheses and associated test statistics from behavioral science papers. The best public eval of this kind of extraction I’m aware of comes from the recent SCORE project - they had humans code thousands of psychology papers to extract their claims by hand. It would be extremely helpful to the world if all scientific PDFs were available as structured open data. I’ve been working to make this happen, both directly at Berkeley and through coordination with large entities I can’t yet speak of; as hard as it is to do, I think it’s possible2. Forensic audit A lot of work has been done on forensic audit, but some gaps remain. Of course, for biology papers that rely on images for evidence, there are a variety of tools (notably Proofig and ImageTwin) to spot anomalies. These are still well short of what sleuths like Elizabeth Bik can do on her own, but these tools are constant companions among fraud analysts. There’s someone working on auditing Excel files for anomalies, and a number of teams are automating numerical checks like GRIM and SPRITE, including Scrutiny project, the INSPECT-SR team as well as statcheck. The regcheck team is building a way to use AI to compare preregistrations to analyses in papers, to ensure there aren’t significant deviations. Nonetheless, there are many other kinds of anomalies to screen for, both public and less publicly known. And there are no formal evals for anomaly detection that I’m aware of. Still, there’s a lot to draw from in this space and I’m pretty certain we will be able to scan papers for most kinds of obvious anomalies in the near future. Methodological review This area has been white hot, though I fear for many of the startups in this space, because this capability may become commoditized. There are at least six different AI peer review companies, including Refine.ink, Reviewer3, ReviewerZero.ai, Q.E.D. Science, Paper Wizard, and Isitcredible. Coarse (a pun on refine) was also recently created as an open source alternative. These systems provide qualitative feedback on the content of papers, spotting methodological weaknesses and mathematical errors. They seem to work pretty well, and many academics report bitterly that they exceed the average quality of typical peer reviewers. But there are few evals here either. What evals exist so far involve using LLM-as-judge (circularity problems abound) or comparing against human reviews of questionable quality. What you’d ideally want is an eval that measures capturing known errors in papers3. Reproducibility and robustness Another active area has been using AI agents to automate computational reproducibility4 and robustness5 checks in papers that report numerical results. For more recent papers where data and code are available, AI agents can see whether they can re-run the analyses and produce the numbers reported in the published paper. In addition to a handful of individual academics who have been experimenting using Claude Code for this, the Institute for Replication is a leading group working on building an end to end system. The evals related to this problem are the most mature, with CORE-Bench (Siegel, Kapoor, Narayanan 2024) and PaperBench (OpenAI 2025) available to benchmark agents on this task. There is also work on getting AI agents to test alternative ways of analyzing the data to ensure the results are robust to small analytic design choices. Synthesis This is the most underdeveloped area where significant investment is required. Although some automated evidence synthesis systems exist — for example, otto-sr is building an AI agent to write systematic reviews — none of these incorporate the full range of paper level signals to weight evidence appropriately. Nor is there anything like an eval or a gold standard for a good systematic review. Arguably Cochrane reviews are the closest we have to gold standard human systematic reviews, though I’ve heard academics in the know complain about their uneven quality. A key question for a synthesis platform is how to weight anomalies and methodological issues in assessing the quality of a piece of evidence. This is an unsolved problem and one I’m very keen to work on. Continuous updating There are many pieces of basic infrastructure available for monitoring for new research and initiating updates. OpenAlex is the current open citation graph. Retraction Watch integrated into Crossref in October 2023. Scite tracks how citations support, contrast, or mention prior claims. The Living Evidence Network demonstrated continuous-update workflows in clinical guidelines. Engineering this is a relatively straightforward task. When you look over this technical architecture and all the progress being made, it’s hard not to be optimistic that a living guide to scientific evidence will be built. The stakeholders are ready The social infrastructure for this is starting to coalesce — it’s not just a pie in the sky academic exercise to imagine this coming into existence. Institutions like the Center for Open Science, the Institute for Replication, the INSPECT-SR, the Living Evidence Network and more are all working on scaling work to improve research quality. Funders are also aligned. The Sloan Foundation has funded living evidence work through COS. Coefficient Giving supports the Institute for Replication and COS. The Astera Institute and the Institute for Progress have shown interest in this space. NIH has established an Office for Replication and Reproducibility. Although there are (very unfortunately) serious headwinds in science funding generally, there is an active group of funders interested in metascience. A brief word about what I’ve been doing at RDI. First, as a Visiting Scholar at Berkeley I’ve been actively figuring out how a non-profit and a public university can conduct and make public the results of large-scale academic article data mining. With some of the money I raised from donors, I commissioned a legal analysis of recent case law and publisher text data mining (TDM) agreements in order to understand whether a massive open data mining of academic articles is possible (with caveats, it is). I’ve also been working to bring together stakeholders in this space, and identify gaps. I’ve also been doing some software development in this space, with more to come. A pilot proposal The assumption undergirding all of this is that an AI, given all this information, would make the right judgment about a scientific claim with lots of conflicting evidence, weighing all the factors appropriately. That’s the hypothesis we need to test. Randomized control trial research is the best place to focus on first. RCTs are used to make many of the important decisions in society - from medical trials to public policy changes. And they use a relatively uniform set of inferential statistics with lots of known and available diagnostics. Behavioral science experiments, within the broader realm of RCTs, should be first used as a testbed whose results can be generalized. Because behavioral science is at the vanguard of open science practices, replications abound (there are thousands of them) to serve as ground truth training data. Key Hypothesis Therefore the pilot would test, in behavioral RCTs that have been replicated, whether the quality of evidence for a claim can be used to accurately predict whether that claim will replicate. Secondary Hypotheses Compared to claims that replicated, non-replicated claims demonstrate a greater share of forensic anomalies in their source literatures. Hypothesis level claim and statistic extraction is accurate enough to scale living evidence without onerous human review costs. Replication prediction is more accurate than prediction markets6 or journal prestige. If all the different quality signals we gather do accurately predict which studies will replicate, then we can use that model to score evidence to power the living evidence layer. Why this is informative regardless of outcome If the pilot succeeds, the architecture extends to medical RCTs (where Living Evidence already operates and integration is mostly about claim representation), then to slices of basic biology with stable replication structure. If it fails, the field learns which quality signals are load-bearing and which ones metascience has oversold. Either result is a contribution to knowing what the literature supports. Implications for funders Because this burgeoning ecosystem of builders already exists, a well-informed philanthropic or government funder could play a crucial catalyzing role in bringing this future about. They could play at least three roles: creating open structured datasets, publishing open benchmarks and incentivizing the solving of key technical challenges. First, open archives of papers that the government maintains, like PubMed, could be turned into structured data amenable to large scale metascience and claim aggregation. I know the US government already has an interest in doing this, though some key open questions remain unanswered. How do you determine the best models and systems for accurately extracting information from papers? How can you establish a robust way to allow researchers to flag errors and correct them? And finally, how do you create a legal regime in working with publishers to maximize the scope of available papers? For the latter, a university or a private philanthropy may be better positioned to make structured data publicly available under journal subscription terms or fair use. Philanthropists or government funders could also coordinate to create or commission benchmarks that evaluate whether important problems have been solved. For example, an open benchmark for claim extraction from a range of different scientific article types would be extremely helpful. Ensuring the underlying data are accurate is vital for creating these evaluations. I’ve discussed opportunistically using various human-created datasets for this purpose, but a consistent problem is that errors in human data make it difficult for them to serve as a gold standard. Finally, with structured data and benchmarks available, the government or private philanthropy could use them to incentivize groups to develop machine learning systems that meet these benchmarks. Prizes are one potentially valuable tool for this. For example, you could establish a prize for a system that accurately updates a living meta-analysis for a small set of claims. Prizes are particularly useful as signals of problem importance, and can help create vibrant ecosystems of public and private research—see, for example the role the government played in kickstarting the current work on nuclear fusion. Conclusion The drawbacks of the current scientific publishing system are known. Scientists agree, metascientists agree, philanthropists agree: the published PDF plus citation graph isn’t the right substrate for maintaining a representation of the evidence base in science. The pieces needed to build the alternative either already exist or are rapidly taking shape. The community is forming around exactly this problem, with concrete partnerships and shared infrastructure. A pilot should start on behavioral science RCTs because that’s the slice of empirical science most amenable to legibility, where replication ground truth is richest, and where the failure modes are best documented. What’s been missing is the galvanizing mission to assemble these pieces into something that works. That’s what I’m proposing to build. *** 1 I have a broader field map that I’ll release publicly soon. This is me, building in public! 2 If this is something that you are excited about, please reach out and talk to me. 3 More on this very soon too! 4 This tests whether, given the code and the data, you can get the same statistics as reported in the published paper. 5 This tests whether the results are the same as a paper’s given alternative analytical decisions in conducting the analysis (like outlier omission). Closely related is the idea of a “multiverse” where you come up with many different ways of answering the same underlying research question with the same data, and test whether the results hold in all those alternative methods. There’s been work on the latter as well. 6 Some of the replications, e.g. those from the SCORE project, had paired the experiments with forecasts from prediction markets. So we get to look at this for free.
Code Puppy: Walmart's secret weapon against AI lock-in
Walmart's viral Code Puppy AI tool helps avoid vendor lock-in, cut costs, and reduce dependence on Claude Code and Codex.
AI Agents Plunged the Tech World Into Chaos. Here’s Exactly How That Happened
The definitive story of how Claude Code and OpenClaw kicked off computing’s biggest transformation possibly ever.
Inside startups, Claude has already won the AI coding wars.
In a survey of more than two dozen startup founders and VCs, we found a growing consensus that Claude Code has become the dominant AI coding tool.
Introducing Shortcuts Playground: Create Apple Shortcuts with Claude Code or Codex
Today, I’m pleased to introduce something I’ve been working on for the past six months: Shortcuts Playground, a plugin for Claude Code and Codex that can create any shortcut for Apple’s Shortcuts app using natural language. With Shortcuts Playground, you can simply prompt Claude Code or Codex with a sentence requesting a shortcut of any
Google Antigravity 2.0 goes after Claude Code and OpenAI Codex with a full agent-first rebuild
Google has taken the IDE off Antigravity. What is left, announced at I/O 2026, is a standalone desktop app built around agents rather than a code editor with AI bolted on.
Microsoft starts canceling Claude Code licenses
Thousands of Microsoft developers will use GitHub Copilot CLI instead
Clawdmeter – A DIY ESP32-S3 desk dashboard for Claude Code token usage monitoring
Clawdmeter is a DIY ESP32-S3-powered desk dashboard that displays Claude Code token usage on a 2.16-inch AMOLED screen so you know when you're about to
Airbnb's CEO says AI writes 60% of the company's code
CEO Brian Chesky added that managers are also getting their hands dirty with coding or using Claude Code.
Anthropic's Claude Managed Agents can now "dream," sort of
Also, rate limits will double for Pro and Max users of tools like Claude Code.
After pushback, Amazon rolls out Claude Code, Codex to all employees
Some Amazon staff had complained about a lack of access to top AI coding tools, arguing the company risked falling behind in developer productivity.
So long Jeeves and Ask.com, relics of yesterday’s internet
Before Claude Code, Grok and Gemini, Jeeves was there in a modest, everyman suit. We conversed with him in full sentences and asked him whole questions. No more.
If Claude Code is going away for Pro users, I can't recommend Claude anymore
This one hurts to write.
Anthropic doubles estimate for Claude Code token spend
Anthropic now says that the average cost per developer per active day is $13, up from $6 earlier this month.
Anthropic acknowledges Claude Code issues, denies 'nerfing'
Anthropic said it found three issues with Claude Code after users complained the AI tool deteriorated.
Ranked: AI Models U.S. Businesses Pay For
OpenAI has long been the leader for paid usage by U.S. businesses, but Anthropic has closed the gap with tools like Claude Code and Cowork.
OpenAI’s big Codex update is a direct shot at Anthropic’s Claude Code
Codex can now use your macOS apps on its own.
We tested Anthropic’s redesigned Claude Code desktop app and 'Routines' -- here's what enterprises should know
For the enterprise, the Desktop GUI is likely to become the standard for management and review, while the CLI remains the tool for execution.
Adobe takes Creative Cloud into Claude Code-esque territory
This is a big step in a new strategic direction for Adobe.
Anthropic closes in on OpenAI as US business use surges
Divergence reflects company’s recent rapid growth owing to strong interest in its Claude Code products
Cursor Launches a New AI Agent Experience to Take on Claude Code and Codex
As Cursor launches the next generation of its product, the AI coding startup has to compete with OpenAI and Anthropic more directly than ever.
Claude Code can now take over your computer to complete tasks
But Anthropic urges caution as "research preview" safeguards "aren't absolute."
Anthropic’s Claude Code and Cowork can control your computer
Its only a research preview for now.
Anthropic just shipped an OpenClaw killer called Claude Code Channels, letting you message it over Telegram and Discord
The consensus among early adopters is that Anthropic has successfully internalized the most desirable features of the open-source movement—multi-channel support and long-term memory
I tried a Claude Code rival that's local, open source, and completely free
I was curious if Block's Goose agent, paired with Ollama and the Qwen3-coder model, could really replace Claude Code. Here's how it worked.
Can AI ‘vibe research’ replace social science?
AI tool improvement is compounding fast enough for researchers to start using tools like Claude Code for real social science tasks.
I tried a Claude Code alternative that's local, open source, and completely free
I was curious if Block's Goose agent, paired with Ollama and the Qwen3-coder model, could really replace Claude Code. Here's how I got started.
Claude Code is suddenly everywhere inside Microsoft
Microsoft is increasingly adopting Claude Code
Anthropic launches Cowork, a file-managing AI agent that could threaten dozens of startups
The new tool brings Claude Code's capabilities to non-technical users, positioning Anthropic to compete with productivity tools like Microsoft Copilot.
Anthropic’s new Cowork tool offers Claude Code without the code
Built into the Claude Desktop app, Cowork lets users designate a specific folder where Claude can read or modify files, with further instructions given through the standard chat interface.
The creator of Anthropic's Claude Code likes to hire engineers who do 'side quests' like making kombucha
Claude Code creator Boris Cherny said that he looked for engineers with "cool weekend projects." Anthropic is also looking for "generalists," he said.
Claude Code's creator explains the limits of vibe coding
The engineer behind Claude Code says vibe coding works for prototypes, but today's AI models still fall short for maintainable software.
Claude Code is coming to Slack, and that's a bigger deal than it sounds
Anthropic launches Claude Code in Slack, letting developers delegate coding tasks from chat threads. It's part of a shift toward AI-embedded collaboration that could reshape software workflows.