Posts
When agents can be highly customized and continuously self-optimize, they become a company’s real moat.
#AIAgents #AgenticAI #SelfEvolvingAI #CompetitiveAdvantage #AIMoat #EnterpriseAI #AIStrategy #FutureOfWork #DeepSeek #AgentHarness
Some principles for building consumer AI products:
- Keep it small. Focus on solving one clear problem. Skip the big vision.
- Start from your own real needs. Only build things that solve problems you’ve actually experienced in daily life or work.
- It’s fine if similar products already exist — but yours must have one distinct angle that makes it meaningfully different.
- If the product involves AI, it must support BYOK (Bring Your Own Key) or integrate a local agent. No forced dependency on a single provider.
- Don’t make free products. Keep the price low, pricing transparent, and keep shipping updates.
- No community groups. Handle communication through email + issues. Continuously improve the documentation and make it AI-friendly.
- Offer easy refunds. Don’t waste time and energy on low-quality customers — just refund them and move on. Your time and mental health matter more.
#AI #ArtificialIntelligence #IndieHacker #BuildInPublic #SoloFounder #OPC #Bootstrapped #ConsumerAI #BYOK #LocalAI #ProductThinking #SaaS
“What’s your team’s competitive advantage if agents can do everything?”
Agents haven’t taken humans out of the loop. They’ve changed what humans do in the loop: from execution to setting direction, managing effective agent work, and judging the output. As execution becomes commoditized, human judgment becomes the real differentiator.
If you doubt the power of philosophy, try this: take a piece of AI-generated text and ask an AI to perform an adversarial review of it through the lens of Aristotle’s final cause—in other words, examine whether the text actually fulfills the purpose for which it exists.
The result may surprise you.
With AI, when building an app, you should spend 60% of your time building and running its marketing system. Otherwise, overall, you’re doing it wrong.
Imagine every worker ant in a colony suddenly became 4x more productive.
Soon, they’d start hauling in everything they could find, building whatever they wanted, and ignoring the original design.
The colony would quickly be DDoS’d by redundant rooms, bad materials, and endless expansion.
Eventually, the ants would have to stop building and start repairing a nest collapsing under its own weight.
That’s what AI is doing to software development teams in 2026.
#AI #SoftwareEngineering #AICoding #EngineeringLeadership #DeveloperProductivity #TechnicalDebt
The most important skill of a great QA engineer isn’t automation, scripting, or designing test plans. It’s an obsessive instinct for spotting anything that feels off.
If you adopt a mathematical perspective, neural networks are nothing more than manifold mappings and optimization in high-dimensional space.
Whether word vectors or latent spaces, the core way AI processes data is by mapping real-world concepts into high-dimensional continuous spaces.
Gradient descent is not merely something in the code—it is the descent of potential energy on a high-dimensional non-convex surface. Where are the minima of the loss function? Does the manifold collapse? Once you master measure theory and differential geometry, you will understand why models exhibit remarkable generalization ability along certain dimensions.
Gradient descent is not merely something in the code—it is the descent of potential energy on a high-dimensional non-convex surface. Where are the minima of the loss function? Does the manifold collapse? Once you master measure theory and differential geometry, you will understand why models exhibit remarkable generalization ability along certain dimensions.
#ArtificialIntelligence #MachineLearning #DeepLearning #NeuralNetworks #Mathematics #DataScience
#SampleBeatsReport
Checking actual samples beats reviewing reports.
I’ll write an article explaining why when I have the time.
I’ll write an article explaining why when I have the time.
How do you push back against a trend?
Sometimes the smartest move is to fully lean into it first.
Sometimes the smartest move is to fully lean into it first.
Only after I’ve actually let AI run unattended for ten hours and deliver a genuinely useful, polished app will I feel I’ve earned the right to say things like “bring the human back into the loop.”
Until then, people will just say I’m not AI-native enough.
If your team isn’t building a content farm or an app farm, don’t chase highly automated, unattended development. Automation is a natural byproduct of mature processes—not something you can force by treating automation itself as the goal and expecting your team’s processes to magically mature as a result. Just look at #OpenClaw’s sharp drop in traffic: you’ll see that the hype around spinning up dozens of agents to develop in parallel simply doesn’t work. It’s mostly commercial marketing.
In the process of project development, recognizing the strengths and weaknesses in the abilities of oneself and each team member, as well as their personality traits, then assigning them to roles that maximize the team’s overall output, designing systems to compensate for their common problems, and helping them improve. This entire process itself is a form of system design—supposedly what is known as #TeamTopologies.
In an AI-enabled workplace, the most valuable people do two things well: they use AI to develop the best approach, and they apply strong judgment to review, challenge, and improve its output.
If someone still relies on others to define the approach, while not being rigorous enough in reviewing AI-generated work, that is an area where they need to grow.
Marketing ability far exceeds the ability to make a good product.
For example, my development harness is better designed than several popular ones on the market, but nobody knows about it.
Rather than making a product, it’s better first to make a product for marketing products.
For example, my development harness is better designed than several popular ones on the market, but nobody knows about it.
Rather than making a product, it’s better first to make a product for marketing products.
If a harness system has both ADR (Architecture Decision Records) and a ticket tracking system, it suggests that the system designer has failed to establish a true SSOT (Single Source of Truth).
ADR can be generated directly from the ticketing system. Maintaining two parallel records of the same class of information is not a convenience; it creates the possibility of divergence and leaves two competing sources of truth. Over time, this inconsistency is highly likely to introduce development errors, especially in AI coding workflows.
My development harness Dimension Leaper is better than #MattPocock’s, which is widely regarded as one of the best on the market.
One of the key differences is that #dimleaper supports parallel development across multiple worktrees by default, with each worktree automatically assigned its own unique ports.
That means you can kick off 10 development tasks at the same time. When they’re done, all 10 can already be running on their own ports, ready for you to review.
Ontology is relational modeling dressed in graph vocabulary.
My intuition is that the right way to apply LLMs to large-scale business datasets is to let the LLM design a relational database and then operate on that database, rather than relying on an excessively complex graph database. But Palantir’s heavily promoted Ontology model appears, at first glance, to be a knowledge graph. So I looked into it.
What I found is that it looks like a graph, but it does not have a graph engine. Relationships between objects are indeed first-class citizens, and the modeling process feels like drawing a network of connected entities. Underneath, however, many-to-many relationships are stored through join tables, creating an object type requires selecting a column as its primary key, and traversal is performed one hop at a time in code.
In Ontology modelling, there is no graph query language, and cross-Ontology joins are not supported. In other words, it behaves very much like a relational database. It adopts the vocabulary and presentation of a knowledge graph, while its implementation remains firmly relational.
In a world where AI can do almost anything, it is already clear what kind of employees are needed: people who can independently solve open-ended problems to a high standard. If the solution is already known, the problem is better left to AI.
Everyone has intentions, yet few can collaborate with frontier AI to create. If what is most precious about humans is not intelligence but each person’s unique original mode of thinking, then collecting intentions from the masses becomes meaningful—and when intelligence is available on demand, this grows even more significant.
Looking back at the GPT-5.5 days, I was especially careful when designing the 3rd generation dev harness. Take the grill-me skill, for example: I carefully figured out what it should be prompted to grill, whether it should grill requirement, spec or plan, where the grilled output should be stored, and where the grilling decisions should be saved.
By the time of the 4th generation dev harness—when we already had GPT-5.6—I simply told the AI: “Add a grill step after the plan in the development process, then write the results back into the plan.” It works much better and faster than 3rd harness on GPT-5.6.
The most important ability for developers today is the ability to solve open-ended problems, not the ability to execute according to a ready-made plan. Because once an executable plan exists, it’s better to just let the AI do it—the AI is far superior in both speed and quality to a mere intermediary.
Fun fact: even when you use an AI as capable as Claude Opus 5 high to write a skill containing around 2,000 lines of text, it can still contain many bugs. It needs at least two rounds of professional review before it is ready for production.
When you think your skill has no bugs, it is often because the runtime AI is silently compensating for them.
The output is free. The cost is paid by those who review it.
That’s my biggest takeaway from the flood of AI-generated code.
Advice for developers:
- Read requirements line by line — AI can rephrase, but you must still read them yourself.
- Align on a clear plan with AI first. Fully understand it before coding.
- Personally verify results against acceptance criteria, one by one.
Advice for requirement owners:
- Write the requirements yourself, or review every line carefully.
- Never just pass an AI-generated document to developers.
My development harness has now reached its 4th generation.
The 1st version, T2P, ticket to PR, automatically created a worktree from a ticket, reviewed the code, and opened a pull request.
The 2nd, IntentMill, could turn a ticket into a spec, grill-me, plan, implementation, tests, and pull request.
The 3rd, ProdFarm, could create its own tickets, develop them through IntentMill, and deploy the results.
The 4th is DimLeaper—short for thinking beyond LLMs and across dimensions. Built around the capabilities of the latest AI models, it radically simplifies ProdFarm and runs roughly 10x faster.
It also corrects an overly idealised assumption in the earlier system: that completing a ticket was the end of the process. In DimLeaper, human review and rework after implementation are first-class parts of the workflow.
Here is the prompt I use to let Claude take the initiative and work with continuous high quality:
Get things cleaned up nicely, and if something doesn't work out, open the tickets yourself to fix it. Keep working until you can say, "I, Claude Code, am basically satisfied now." Before you can say this sentence, it is forbidden to stop and ask me to make decisions. I'll give you a budget of 10 tickets. Start now.
Then I saw it spin up 20 subagents to investigate and fix every issue it found. I realized I couldn't design a development workflow like this for it that would yield the same results—it is time to let go.
Noticed something interesting: AI often doesn't trust that letting itself work autonomously gives better results. It always wants to write rigid, mechanical code to handle a problem, rather than spinning up a subagent and directing it to explore with a simple prompt. So the thing holding AI back from its full potential is quite possibly AI itself.
Maintaining 1,000 E2E tests is a full-time job. Maintaining 1,000 two-sentence test points isn't. That's the shift my team landed on over the past few weeks—less a tooling change than a change in what we ask AI to be.
Your Decomposition Is the Bottleneck
I kept asking AI to solve a problem. Every answer disappointed me. Then I stopped describing the problem and described the outcome I actually wanted.
It saw immediately what I hadn't: the problem only existed because of a path I'd assumed was necessary. The goal didn't require solving it at all. AI took a different route — and it worked.
My mistake was defining the wrong search space. I'd decomposed the task myself, then confined AI to optimising one intermediate step, never letting it ask whether that step belonged there.
State the outcome. Let AI find the path. Your own decomposition may be the very thing standing between you and the answer.
The capability of any developer — human or agent — is the ratio between the work of theirs you accept and the instruction it took to get it. The higher the ratio, the more capable they are.
A person must pursue perfection within limited resources and be willing to obsess over every detail; otherwise, they cannot be entrusted with major responsibilities.
While working on our product’s UI, I found that the secondary colour did not integrate well with the rest of the interface. After looking into it, I realised that the traditional concept of a secondary brand colour has largely fallen out of favour in modern UI design.
A growing problem with AI-assisted development is that it makes it easier to avoid the real work.
We read summaries instead of design documents, test reports instead of using the product, and conclusions instead of looking at real data samples.
But once we stop working with the details, we also lose the ability to ask good questions.
We should use AI to get closer to reality, not further away from it.
This prompt keeps Claude Code running continuously, creating value while I work or rest: Use playwright-cli to inspect [URL] yourself, identify the main issues, and fix them directly. You can use the Plane skill to create tickets for tracking. Fix one issue, deploy it, check the site again, and then continue with the next one. You are authorised to create up to 10 tickets. Prioritise the most important issues, with customer acquisition, retention, product stability, and the elimination of obvious errors as the main goals.
I’m getting a bit tired of all these Loop Engineering and Graph Engineering labels, so I decided to invent one more: Space Engineering — one concept to cover them all and hopefully end the naming game.
The hardest parts of software development have now become feature planning, product structure, copywriting, and UI design. Once those are settled, the software itself can almost be generated automatically.
Good design doesn't come from inspiration. Here's the full four-step process I use to build website UIs that are genuinely polished and professional.
Codex’s Goal mode has a serious bug that I have encountered three times. While executing a goal, it may unexpectedly rerun commands that were executed before the goal started.
For example, I first asked Codex to delete a document, then used Goal mode to spend two hours generating a new one. During the process, it suddenly reran the earlier deletion command and deleted the newly generated document. It then fell into an infinite loop, burning tokens while producing nothing, and potentially even damaging the system.
As of July 18, among all OpenClaw PRs:
- Fix PRs account for 62%
- Feature iteration-related PRs only account for 8.5%
- Refactoring-related PRs account for 26%
It looks busy on the surface, but most of the parallel tasks are just patching and fixing small issues. These tasks don’t require humans to provide much context, so the system can run a lot of them at once.
Moreover, it doesn’t look very stable. For example, PR #78595 was a SQLite migration that changed 3,135 files. It was completely rolled back just 18 minutes later (commit 694ca50e), and then followed by another 138 fix commits.
This is a classic example of AI slop. In a typical commercial company, this would be considered a P0 incident.
Here is a small example that shows why human-in-the-loop is still necessary in agent coding.
I built a small Language-learning app. In the dialogue section, all speech is generated automatically with TTS. During testing, however, I found that Claude Code with Fable 5 had assigned a female voice to a male character and a male voice to a female character. It had correctly realised that the two speakers should sound different, which was already a non-trivial detail, but it still failed to match the voices to the gender implied by the characters’ names.
This illustrates a broader limitation. AI can still miss very small but important details, and often cannot even identify them during the grilling stage and ask the human for clarification.
You might argue that I should simply use a popular specialist TTS skill. But integrating and operating another specialised skill would require more time and effort than the issue justifies.
At this stage, AI still needs a human at the end of the loop to review the final output.
So before boasting that you can run dozens of development tasks in parallel on one project, first ask yourself whether you have carried out a thorough human acceptance review as much as your paralleled tasks.
Claude Chat has far stronger product thinking than Claude Code, even when both use Opus 4.8 High. Make sure to distinguish between them.
My ticket template contains only three sections: scope, constraints, and acceptance criteria. Every one of them exists to establish boundaries.
Defining boundaries is a far more important capability than solving problems. Once a problem’s boundaries are clearly defined, the solution will emerge almost automatically by AI —and will often be better than anything a human could plan or design.
How powerful is a real harness once it starts running? Assuming a 10-hour workday, one project can complete around 20 tickets with generally satisfactory delivery quality and only about four rounds of human instruction.
Based on my current estimate, one person can run five projects in parallel—meaning up to 100 tickets delivered per day. These are not artificially tiny tickets; Claude breaks them down according to software engineering best practices.
But operating at this scale requires abandoning perfectionism. You cannot obsess over every word, interaction, or pixel. You have to focus on what materially affects the outcome and ignore the rest. Each product should be treated as an ROI machine, not as a handcrafted object that must be perfect in every detail.
#HarnessEngineering #LoopEngineering #AI #Agents #Codex
The secret to making an agent work methodically is the output contract: ask for a step-by-step execution log, and it must pay attention to every step; ask only for task completion, and it will tunnel-vision its way to the finish line.
I found the quantifiable objective function of an AI harness: minimise the number of human decisions required during execution while also minimising the number of rework points requested after delivery.
I built an AI persona skill to let a panel of simulated users test an app before real users do. Here’s how it works! #AgentSkills #LoopEngineering #HarnessEngineering
With AI, companies do not need large numbers of employees. They need two kinds of people: genuinely intelligent employees and highly accountable employees. Intelligent employees can use AI to increase their productivity by 100×. Accountable employees ensure that every deliverable undergoes human review.
Design constraints are first-class citizens in a harness system, which consists of two main categories: SOPs and design constraints. SOPs ensure that the overall workflow does not create chaos. Design constraints ensure that implementation details remain aligned with human intent.
Design constraints include the complete UI design system, data ETL, data model design, technology stack, coding standards, test users and testing methods, and red lines that require explicit human approval.
Consequently, at the solution design, code development, and ticket acceptance stages, these constraints must be treated with the exact same priority as the requirements themselves.
The core principle of my harness system is that the user provides a research proposal, prototype, or initial idea as a seed. The harness then conducts research on the codebase, market, tools, and data sources, and converts that seed into no more than ten tickets managed in Plane.so.
Each ticket must go through several stages: draft, specification, development plan, implementation and testing, acceptance, merge, and deployment.
Using this approach, I can keep Claude Code or Codex working continuously for more than five hours, delivering highly complex feature sets or changes with generally satisfactory results. Everything remains traceable through well-structured tickets.
GPT-5.6 is revolutionary.
In the past, GPT was worse than humans at filling in unstated goals and constraints. Now, it can often do this better than humans.
So with GPT-5.6, making fewer decisions can actually help it build better products.
How vibrant is the AI startup scene? Hundreds of new products, tools, and business models emerge every day. It feels like living through a nuclear explosion on a daily basis.
A skill that prevents your prompt from making the same mistake 200 times
Yesterday, I asked an agent to process more than 200 items. The first output was wrong, but it kept using the same prompt and repeated the same mistake across the entire batch.
So I designed an automated closed-loop batch skill:
Process a small batch → validate the results → fix the outputs → automatically update the prompt → retry → continue with the next batch.
Once several batches pass consistently, it switches to sampling-based validation.
The goal is simple: do not discover 200 errors at the end. Make the agent learn automatically from the first batch.
I have built a harness that lets me state a product goal in a single sentence and have AI autonomously spend several hours building and operating an app. The question is how to make real money from these one-sentence apps.
A friend used to build one app a month that nobody used. Now he spends $3,000 a month on tokens and builds 37 apps that nobody uses.
Where is the breakthrough?
Answer: Compared with building apps, market research and marketing matter more. Then comes using AI to find online arbitrage opportunities or run content farms. That is often more profitable than building apps.
Since today is the last day Fable 5 is free, I have to finish building ProdFarm — a skill that can automatically launch a product from product intent.
EvoDocs, the skill I created, automatically generates product documentation from code.
Recently I realized that documentation also has a hierarchy. The most important documents are the product-level documents: product goal, feature scope, tech stack, architecture, data model, and operational runbook.
These matter more than documents explaining how each module works, or how modules interact internally.
I think people with these two qualities can maximize Fable 5’s capability.
First, the philosopher-like ability to abstract: the ability to simplify and recombine an endless stream of tools at will.
Second, a kind of imperfect, almost reckless creativity: listen to AI, make only a small number of decisions, and do not let yourself become the bottleneck that slows AI down.
In AI coding, who would have thought that the most necessary and hardest thing for humans to get right is file management?
More specifically: how you organize your code, knowledge base, skills, configurations, infrastructure, and all online and offline files and directories.
A team that is messy in these areas will inevitably produce messy output.
I spent half an hour building an EasyApp skill that can deploy any full-stack web app to Azure with a single command. Each app gets its own backend service, database, and CSR/SSR frontend, without adding any extra cloud infrastructure cost. This skill allows our team to launch any internal tool or product prototype with essentially zero cost and zero setup time.
I propose Midday Check-in for agent coding teams. Daily standup should review yesterday’s outcomes and align today’s plan. Midday check-in is not about pushing progress, but about preventing deviation. When AI greatly increases output speed, a wrong direction can quickly turn into hours of bad code.
LangChain’s newly launched #OpenWiki is almost identical to #EvoDocs, which I created a few months ago. Both generate documentation from code and automatically update it based on commits.
All tools will eventually be replaced. The only thing that will not become obsolete is the money made by products built with those tools.
After seeing the power of loop coding, I have a few reflections.
The data layer may no longer need to be carefully handcrafted. What LLM needs is a messy cloud of data: for example, all historical information about a product on Azure. It can freely mine that data, surface patterns, and visualize what matters. Humans do not need to manage the data sources, data models, or data pipelines in the traditional way. Let the LLM explore.
Software can naturally grow out of chaotic data.
Software is essentially a machine that turns data into value. Its shape is determined by two things: the data it has, and the value it is meant to deliver. Given those two forces, the shape of the software may not be arbitrary at all. It may even become inevitable.
And this machine can now be generated automatically by an LLM within a few hours.
I am using loop engineering to build a brand new Product Operating System.
The idea is simple: give Codex only the product objective, then let it use my harness skills and Plane.so to create its own tickets, implement them, deploy them, and repeat the loop.
After running it for three hours, I was shocked.
The completeness of loop coding is far beyond what vibe coding can achieve.
My reflection is that in vibe coding, humans spend too much time obsessing over product details and implementation plans. We try to replace AI with human intelligence, and the result is lower efficiency and lower quality.
Loop coding feels different: define the objective, build the harness, and let the system continuously move toward the product goal.
After automating tickets from requirements to delivery, my next goal is to automate the entire product lifecycle from objective to delivery. The goal after that is to automate product growth.
A quantifiable way to measure how AI-native a developer is is to look at how many agent skills they have created proactively over the past six months.
The most important personal quality in the AI age is agency. As long as a person has agency, they can use AI to achieve a 100x productivity gain. But if someone is only satisfied with doing assigned work, they will only get a 4x improvement.
The secret to one person creating and managing dozens of projects is not to pursue perfection. Once the product design is decided, hand it over to Codex. Ignore the details as long as there are no major deviations.
Second, establish #LoopEngineering to cover the whole loop from creating tickets, development, testing, deployment, maintenance, operations, and creating tickets again. This entire loop is controlled by mature #HarnessEngineering and automatically completed by Codex. Humans check the output once per day.
This is not impossible.
How can AI generate and iterate projects based on intents? The following steps should be fully automated:
- First, create the GitHub repo, server, domain, database, LLM, and other resources.
- Initialize the project framework: TanStack Start, Python pipelines, Prefect, crawl4ai, Clerk, PostHog.
- Initialize my harness system: EvoDocs, IntentMill, AutoQA, and plane.so, Codex/CC, design.md.
- Based on the project goal, have Claude Code write requirement tickets in plane.so, develop with the harness, validate the work, and use this as the loop. Humans review the output once per day.
Data-driven capabilities and harness engineering are the moat of IT companies in the AI era.
Why can people reverse-engineer Open Design from Claude Design, but not reverse-engineer Claude Code? Because outsiders can only see the UI surface, not the large amount of system logic behind it, especially the data-driven product capabilities built on massive amounts of data and user behavior data.
You can only copy the existing visible features, but you cannot copy the engineering delivery capability that allows Claude Code to ship stable new features quickly.
The most informative takeaway from ByteDance's FORCE conference wasn't the product name TRAE Work, but a set of real internal metrics: over 90% of the code written by the TRAE team is AI-generated, yet per-engineer feature throughput has increased by only about 60%.
This doesn't mean AI coding has failed. A 60% throughput gain is already substantial for any mature engineering team.
What it really reveals is a different bottleneck: generating code has become cheap, but the pipeline that turns code into production-ready software hasn't kept pace.
It's very similar to what happened with DevOps.
Once Agile dramatically increased the frequency of code changes, traditional integration, testing, and deployment processes became the bottleneck. DevOps and CI/CD weren't optional improvements—they were created to ensure high-frequency changes could reliably make it into production.
ByteDance estimates that agents can deliver a real productivity gain of about 1.6x for programmers. Whether a project can scale and keep iterating sustainably depends on how far its harness engineering can go.
A technical lead must stand firm on the design direction of every project and should not compromise simply because developers disagree. Otherwise, the project will become increasingly difficult to maintain, and in the end, the technical lead will still be responsible for fixing the consequences.
Today, while using the intent-grill harness system for development, the fully automated development and testing flow completed successfully. But during my review, I found 8 issues that still required interactive refactoring. That does not align with the intent-grill principle.
So I added three new requirements to the grill-me skill:
- UI changes must follow “research first, then grill, then user approval.”
- Any external API integration must follow “read the docs first, predict side effects, then get explicit user approval.”
- The grill must simulate what problems unit tests may hit during development.
Now my harness is already close to a closed loop.
The User only needs to say what they want, then answer a few questions to clarify some ambiguous points. That feature will then be fully automatically developed and integrated into the existing product.
[User] only needs to input the requirement.
[Codex] automatically creates a branch.
[Codex] generates a draft solution based on the documentation and code.
[Codex] generates a list of decisions that the user needs to make.
[User] answers these decisions.
[Codex] automatically creates the spec and plan.
[Codex] uses goal mode to automatically develop and do unit testing.
[Codex] updates and runs the regression test suite.
[Codex] updates the documentation.
I’m not optimistic about #Eve, #Vercel’s newly released agent building framework. What people lack today are agent skills that execute strict workflows, not agents that can freely decide which tools to use. Skills are what enable work to be completed strictly according to human intent. The enormous degree of freedom that agents have means they can only be used for exploration.
https://vercel.com/eve
As software becomes an autonomously evolving system driven by agent-based development, the core competitiveness of creators will lie in their ability to optimize the continuous balancing of marketing, feature iteration, and cost control to maximize long-term returns.
Final-Cause Adversarial Review
The secret of adversarial agent review is to ask first: what was this artifact created to achieve?
This is its final cause: the purpose that gives its form meaning.
Judge everything against that purpose. Does the artifact’s current form help the purpose become real, or does it introduce drift, ambiguity, misunderstanding, and miscalibrated degrees of freedom?
Do not enter the semantic forest too early. A single question aimed at the final cause can be stronger than a long list of surface rules.
The final cause is the measure. Everything else must justify itself before it.
Using one sub-agent to #adversarially review the output of another sub-agent is one of the most effective ways to improve LLM output quality.
This is the software factory workflow I designed. Only steps 1, 6, and 12 require human involvement; everything else is fully automated.
They correspond to: submitting the requirement, answering key decision questions, and reviewing the UI.
Roughly 20 skills support this system, and some of the more complex skills contain more than 5 capabilities.
For code changes involving databases, UI, and architecture, it is impossible to remove humans from the loop entirely.
What we can do is automate all development, testing, and integration, leaving only a “grill me” stage where humans answer a few critical decision questions from the AI.
Software development is dramatically reduced to a handful of Q&A decisions.
A production-grade complex skill is often more than just a main workflow description. It is composed of multiple capabilities that are cohesive around the same business scenario. Each capability should have its own independent directory, communicate through artifact documents as inputs and outputs, and maintain clear responsibility boundaries. However, splitting these capabilities into multiple standalone skills introduces additional complexity in maintenance, version synchronization, and context alignment.
This led me to design the cap-gate-loop pattern. Within a single skill, a complex business scenario is organized into multiple stable capabilities. Each capability has its own directory, instructions, scripts, and output boundaries. For high-risk capabilities, a matching gate with the same identifier is added to perform semantic review of the generated artifacts, preventing outputs that are formally correct but ultimately say nothing of substance. When a gate fails, it triggers a rework loop.
For users who are not good at using AI, the product should not try to teach them how to use AI. Instead, it should proactively surface relevant content, let the AI ask targeted questions, and use the user’s responses to continuously improve what gets pushed next.
Once everything can be made agentic, the scope of what programmers can do becomes much larger.
What should we build to handle marketing, the software factory, and team management well?
These used to be jobs for specialists. Now programmers can do them too.
Or put another way: everyone is a programmer.
Or even more precisely: everyone is an agentic workflow builder.
Claude is strongest in critical thinking; GPT is strongest in logical completion; Gemini is strongest in divergent thinking.
Single Source of Truth is far more than avoiding data duplication — it is the structural core that determines whether a system can stay consistent and evolve sustainably. This article explores good versus bad abstractions in software design through the lens of category theory’s universal properties.
#SoftwareArchitecture #DomainDrivenDesign #SoftwareDesign #CategoryTheory #AbstractionPrinciples
This weekend I spent many hours studying category theory. After I finished learning universal properties, I realized that the four software engineering principles I have always followed are not merely my personal preferences. They are, in fact, the engineering embodiment of the philosophy of category theory:
There should be only one way to do a thing.
There should be a single source of truth.
The design should be minimal and only as necessary as it needs to be.
Fallbacks should not be allowed.
Only by doing this can we preserve the purity of the architecture’s topology, eliminate a large class of potential bugs, and make the entire system easier to understand and reason about.
I ran into an interesting problem: I need to add a capability inside one skill whose purpose is to generate another skill specialized for a specific project. In other words, it is a skill for creating project-specific skills. This is where the bootstrapping process reaches a project-specific fixed point, because the equivalence relation that defines the semantic quotient space cannot be fully generic; it has to be fixed by the domain knowledge of the specific business.
The previous basic programming model was: operating system → application → functions. The current basic model is: coding agent → skill → skill capabilities. These two structures are isomorphic.
Don't be misled by AI influencers claiming that humans are no longer involved in development.
They're only showing you the tip of the iceberg.
What lies beneath the surface could be a support system built over several months, or it could be AI-generated slop riddled with endless bugs. You'll never know.
If you don't understand what's underneath the iceberg, you're better off sticking with vibe coding—it will give you a much better chance of maintaining code quality.
Over the past two months, I have been building and evolving a comprehensive Harness Engineering ecosystem focused on automating the software development lifecycle.
The core initiatives include:
• T2P (Ticket-to-PR): An automated development, code review, and quality gate system that transforms tickets into production-ready pull requests.
• AutoQA: An automated testing framework that validates ticket implementations and continuously expands the regression test suite.
• IntentMill: A solution-generation engine that automatically produces implementation approaches, technical solutions, and test cases for every ticket.
• EvoDocs: A documentation system that generates design and architecture documents directly from source code.
All of these systems are actively under development while simultaneously being used to build and maintain real production code.
In addition, I have developed and maintained a set of supporting tools and skills:
• TicketKeeper: Ticket and sub-ticket creation and management automation.
• ToASkill: A skill-generation framework that produces standardized skills, captures best practices, and helps engineering agents avoid common pitfalls.
• SkillHost: A platform for distributing skills across coding agents and managing skill updates.
• GitPlant: A worktree creation and management system that enables isolated parallel development workflows.
• GetAGoal: A goal-loop framework that helps maintain alignment, iteration, and continuous progress.
Rather than being standalone tools, these components form an interconnected engineering ecosystem. Every system is simultaneously being developed, refined, and evolved, while also being used to accelerate the development of the ecosystem itself, creating a continuous feedback loop of improvement and automation.
Any changes to the UI, database, external APIs, state machines, or prompts are architecture/product decisions. They must not be invented by AI during the implementation phase. They must be explicitly defined by humans at the spec stage.
#grill-me before it hits the harness.
Yesterday I was debugging our harness system.
One requirement was to capture the user’s intent. Anyone on the team would know that this should be extracted from the user’s chat input. But the AI didn’t know that. It decided to create a pop-up asking the user to enter their intent, and it actually implemented it that way.
This shows that most requirements, before entering the harness, must go through a “grill me” step: list out the questions, have a human clarify the ambiguous parts, and make the key decisions.
The structure of agent skills will inevitably become increasingly complex.
For example, a skill may include versioning files, dependency files, multiple feature branches, Markdown/script subdirectories organized by function or module, dedicated test directories, dedicated evaluation directories with the same level of importance as tests, multi-threaded workflows involving multiple sub-agents, temporary working directories, and even its own database.
In short, skills will become the applications of the AI era.
Yesterday I let Codex run on a task for a few hours. Total usage: 8 million tokens. I got curious and looked up the pricing. The comparison below was not what I expected.
Claude's newly launched Workflow is a breakout feature. It embodies proven multi-agent design patterns and provides the essential scaffolding for tackling complex tasks.
Many workflows in our codebase, such as semantic ranking, classification, and batch processing, can benefit from Workflow's sub-agent pattern. This approach can significantly reduce context contamination, improve output consistency, and deliver higher-quality results.
https://lnkd.in/e3zwd9eW
It passed every validator.
It still produced garbage.
In a skill, prompts don’t fix broken workflows.
The goal of harness engineering is to make an agent’s behavior and outputs align as closely as possible with the user’s true intent. To achieve this, we need to design both the agent’s behavioral patterns and its knowledge base. The knowledge base is the most important component, bar none: only after the agent has absorbed the user’s knowledge, background, and working context can it truly align with human intent.
This is the ultimate form of all ToB AI software. If it’s not like this, it is absolutely designed wrong.
Skillhost 0.1.9 is out on PyPI.
This release improves repo consistency handling: if a skill repo is still registered but missing on disk, list now shows it clearly, remove can clean it up without crashing, and relink/update fail with a clear error instead of getting stuck.
https://skillhost.dev/
The essence of software design is abstract algebra; orthogonality and quotient spaces are the most powerful weapons humanity holds in both hands against high-dimensional complexity.
In harness engineering, the most important thing is documentation that updates automatically with the code and the completed requirements. Only a correct and concise documentation set can enable Codex or Claude to generate the right solution and acceptance test cases based on the requirements. Together, the solution and the acceptance test cases determine whether the requirement can be completed automatically.
We're a Codex team. Everyone built useful skills. Nobody had a clean way to share them. That's why I built SkillHost.
pip install skillhostskillhost add <git-repo># clones into ~/.skillhost and symlinks into your agent's skills folderskillhost update# pulls latest, everywhere it's linkedskillhost list# see all skills; mute any without uninstallingCopy code
Add
--project to any command to scope it to a single repo instead of your whole machine. User-level skills follow you everywhere; project-level skills stay put.Works with Codex, Claude Code, and OpenCode, OpenClaw.
No registry, no server — just Git doing what it's good at.
→ skillhost.dev
Shipped a small tool I originally built for myself: SkillHost.
Problem: our team uses Claude Code, Codex, OpenCode — every agent expects skills in a different directory. Sharing skills turned into zip files, stale forks, and constant “which version are you on?” confusion.
SkillHost fixes that with one simple idea:
use a single Git repo as the source of truth for all team skills.
It symlinks the repo into every agent automatically, so the whole team stays in sync. Update once, every agent sees the latest version instantly. Every link is tracked in a manifest, so unlinking is clean and safe.
No accounts. No backend. No platform lock-in.
Just Git + symlinks, done right.
→ https://skillhost.dev
Many code indexing tools have emerged to provide context for AI coding.
But I think they make two fatal mistakes.
First, they should not index code down to the function level. The file should be the smallest unit.
Second, they should not use a knowledge graph. Instead, they should use a connected set of Markdown documents that can be incrementally generated and retrieved by LLMs.
AI can absolutely understand your code. What it lacks is the business process and key decisions behind that code.
Decoupling the M system, which is responsible for product, requirements, and solutions, from the T system, which is responsible for development, testing, and operations, is key to dramatically improving AI efficiency.
Once the M system has completed solution design for multiple requirements, the T system can continue developing in the background. The results are then reviewed and accepted in the M system, which generates new requirements in turn.
This turns automated workflows into automated work loops.
Even while humans are sleeping, AI can continue creating value.
How to Make AI Improve Your Company While You Sleep
In a recent batch talk, YC General Partner @t_blom broke down how to build a self-improving, AI-native company.
He explained how to create recursive, self-improving AI loops, and why founders who get this right will be able to run companies that keep improving while they sleep.
I extracted the following ideas from the video:
- A company is like a Roman legion. The strength of a Roman legion came from the fact that every individual soldier was highly capable. There were no weak links, so the legion would not fall apart when executing any tactical maneuver.
- The copilot model is wrong. You should hand the task fully over to AI, because AI’s execution speed is 10x to 100x yours. If it gets something wrong, just have it redo the work. The second and third attempts will still be much faster than writing the code manually.
- Burn tokens, not headcount. This is exactly right. Any AI-native company should give employees unlimited tokens. Only with unlimited tokens will people be willing to use AI freely, without psychological friction, and fully unleash the creativity that humans are best at.
- Make everything clearly readable by AI. My interpretation of this is: your old documentation was written for humans, and humans usually do not like writing documentation. Now, you should expose as much of your documentation, interfaces, and MCP context as possible to AI. Once you do that, AI can use its organizational capabilities to create entirely new company workflows.
- People are temporary. Context documents are what matter. This is also an excellent point. In the future, the most important assets of a software company will be its requirement documents, context documents, and process documents. Code is easy to recreate, but these documents and the thinking process behind them are the company’s real assets.
Codex just added support for Hooks.
A lot of people still do not know when they should use Hooks.
My view is simple: anything you find yourself repeatedly reminding Codex/Claude to do should probably be taken out of the prompt and moved into a Hook.
For example:
- Format the code after every change
- Run lint / tests before every commit
- Prevent changes to certain directories or config files
- Automatically check for type errors after code generation
- Intercept and ask for confirmation before touching high-risk files
- Automatically inject project context at the start of a session
- Automatically record a change summary when a task is finished
Prompts are for expressing intent.
Hooks are for enforcing rules.
I just discovered a shocking reality.
Linear and Notion can both be accessed through APIs, not just MCP.
I just pulled 69 Linear tickets in only 18 seconds. If I used MCP, it would take 7 minutes and have a much higher error rate.
Notion is the same.
Barely had time to finish Harness Engineering, and now it’s already time to build Agentic Scrum.
Maybe Agentic Scrum is Harness Engineering applied to team-level software delivery.
My engineering philosophy: SOFA 🛋️
•Single source of truth.
•One way to do one thing.
•Fail fast on constraint violations.
•As simple as necessary.
Harness engineering has finally converged into two concrete problems.
First, how to map code into design documentation.
Second, how to map design documentation into a regression test suite.
The better we solve these two mappings, the more software systems can grow automatically.
I used to think building a prototype was easy, and making it production-ready was hard.
Recently I realized it is the opposite: building the prototype is hard because you have to turn ambiguity into the working structure; making it production-ready is mostly a matter of following a systematic process.
One of our agent skills depends on a Notion document as its working standard.
In the past, we were constantly struggling with how to keep that Notion document and the skill in sync. Every time the Notion document was updated, we had to manually update the skill as well.
Today, I suddenly came up with a neat solution: every time the skill runs, it first checks the last updated timestamp of the Notion document. If it differs from the version stored locally by the skill, the skill updates itself before executing.
This cleanly solves the synchronization problem.
But why didn’t we think of this earlier?
After thinking about it, I realized it is because this solution moves up one level of abstraction. Our normal way of thinking is limited to the relationship between the skill and the artifact it produces. But a self-updating skill treats the skill itself as something that can be observed, updated and versioned.
In a sense, it is meta-cognition about meta-cognition. That is why the idea is simple, and yet not obvious.
This also opens up a future direction: skills may also evolve from every execution feedback, with major structural changes requiring human confirmation.
Tool Usage Capability Preference Matrix (Original)
This 2x2 matrix measures individual capability preferences from a tool perspective.
● Horizontal Axis X: From Few Tools left to Many Tools right
● Vertical Axis Y: From Simple Mastery bottom to Deep Mastery top
Quadrant Descriptions
● Top-Left Specialist: Deep expertise in a small set of tools. High proficiency focused mastery.
● Top-Right Versatile Expert T-shaped: Broad exposure to many tools combined with deep mastery in core ones. Adaptable and powerful.
● Bottom-Left Basic User: Limited tools used at a basic level. Sufficient for simple needs but low versatility or depth.
● Bottom-Right Jack of All Trades: Familiar with many tools but only at a surface level. Wide but shallow capability.
This framework helps evaluate personal or team strengths in tool-oriented workflows eg software crafts productivity engineering. Where do you position yourself?
One common mistake when selling AI to traditional businesses: walking in and trying to “reengineer their entire workflow.”
That approach fails almost every time.
A traditional business owner carries the core business process in their own head. It’s the foundation of how they survive and compete. If you come in with an AI system that tries to replace or overturn it, they won’t buy in. Worse, they’ll think you don’t understand their business at all.
The right approach is much simpler: Find a few concrete pain points. Use AI agents to solve those specific problems well. Give them a real sense of value and relief. Once they see tangible results, they’re willing to pay — and they genuinely appreciate the help.
Solve one thing first. Then earn the right to do more.
AI hasn’t reduced programming jobs — it has increased them. Human desire is unlimited; we constantly want things faster, better, and more abundant. If someone thinks AI-driven productivity means we need fewer programmers, they may be underestimating how many new opportunities and possibilities greater leverage can create.
How can we use AI to generate a comprehensive documentation system from a repository with hundreds of thousands to millions of lines of code, such that another AI can use this documentation to develop within the repo and significantly improve ticket implementation accuracy — compared to even a senior engineer using AI directly on the repo?
Someone analyzed all 5000+ accepted papers at ICLR 2026, and it's a good signal who's pushing the research of AI:
China has surpassed the US with 43.7% of the papers
If a team does not have a team-level harness and simply gives AI to each individual developer, the result is usually local efficiency gains but lower global consistency — effectively turning 1 + 1 into 1.5.
At a high level, using AI to maintain a project across a team may sound straightforward.
I recently built a tool for my team that lets AI review code against requirements (ticket to PR), and it already exposed 25 hard design questions.
The hard part is this: a real project is not a single-person closed loop. It is a multi-person collaboration system. One person writes requirements, several people develop, and someone else tests. Everyone’s inputs and outputs are unstable, and not everyone will naturally follow the same rules.
What you are doing is not simply asking AI to write code. You are trying to map unstable human inputs and outputs into an engineering system with hundreds of thousands of lines of code, while keeping the process consistent, auditable, and traceable.
That is the real difficulty of applying harness engineering in a team.
Harness engineering is still engineering. The core question is not whether Codex can write code, but whether the harness can define constraints, detect failures, and safely control code generation. #harness
In the vibe coding era, a five-person team can generate PRs like a twenty-person team before AI.
The bottleneck is no longer code production. It is whether the organization can safely absorb, review, test, document, and close the loop on all those changes.
That is why harness engineering becomes necessary: AI-accelerated teams need closed-loop development management, not just smarter coding agents.
The core of harness engineering is not building smarter individual skills, but validating and handing off the artifacts between them.
Smart coding or review skills alone are not enough. Without a closed loop for validation and handoff, they introduce noise, accumulate anomalies, and eventually destabilize the whole development process.
A reasonably smart model with ordinary skills, running inside a disciplined loop, can outperform smarter models or even skills written by star developers when those skills fail to form a closed loop.
The Codex app is the browser of the AI era.
Its plugin store is the new app store.
Chatting with AI is the modern equivalent of scrolling the web.
A system, or any module within a system, can be described through five dimensions: Objective, Structure, Process, Interface, and Constraints. #OSPIC
Philosophically, these five dimensions can be understood as the system’s ontological categories. From an engineering perspective, they also align with the artifact description framework used in the design, construction, and maintenance of engineering artifacts.
I call this system description framework the OSPIC Framework. OSPIC can be used to structure project documentation.
Harness engineering has converged for me: humans should be forbidden from using Codex to write project code directly.
Humans should only write product requirements — and use Codex to write the code that controls Codex writing the project code.
In vibe coding, humans write objective functions.
In harness engineering, humans write objective functions for objective functions.
I almost used OpenSpec in the ticket-to-PR workflow, but then I realized that we need to respect each developer’s working style and leave room for creativity. So I decided to open up the freedom of design, development, and testing to the user.
T2P will only be responsible for review and archiving: requirement alignment, code review, test coverage checks, and documentation updates.
The ticket-to-PR skill will be TodoClaw’s primary weapon.
When a ticket is assigned to TodoClaw in Linear, it will launch a sandbox and prepare the required project environment and skills to develop that ticket using T2P.
During development, it will refer to the project documentation and strictly follow T2P’s requirements to design, plan tasks, implement, test, review, and then submit a PR.
All human engineers need to do is break tickets down into well-designed, reasonably scoped, and meaningful subtickets; review the final PR delivery; and confirm the merge.
That’s it.
Today I shared a T2P (Ticket to PR) skill I’ve been developing with the team.
T2P is a #harnessengineering automatic workflow that covers the full lifecycle from creating a Linear ticket to completing development, testing, documentation updates, and merging the PR.
Its key characteristics are:
- Ensures every ticket implementation satisfies the stated requirements without drifting out of scope.
- Ensures sufficient test coverage.
- Requires both Codex review and project-specific rules review to pass.
- Requires relevant project documentation to be updated.
- Blocks PR merges if any of the above conditions are not met.
- Preserves maximum freedom during the design and development phase: engineers can work spec-first, use vibe-coding, or delegate the implementation entirely to an agent.
- Captures the development process for each ticket so the work can be retrospectively reviewed and analyzed.
TodoClaw is positioned as an agent runtime platform centered around to-do list tickets, focusing on providing harness engineering around Codex for multiple projects. Its aim is to accelerate the development progress of our company's dev team.
Typical functions include: creating standardized linear tickets, creating standardized documentation, automatically monitoring error logs, fixing bugs and releasing PRs, mandating code reviews for PRs, generating executable front-end and back-end test cases for each ticket, and maintaining regression test suites.
This allows team members to focus on solution design and vibe coding, while todoclaw handles all the tedious tasks and ensures project quality.
I am thinking about integrating #OpenSpec to our agentic development pipeline.
Milla Jovovich (the actress who played the Resident Evil heroine)'s critique of current AI long-term memory systems: faulty memory retrieval and integration. Once these errors gradually accumulate, they make the memory unreliable, ultimately causing new tasks to fail.
This matches my thinking perfectly, which is why I never use long-term memory systems myself.
The solution I proposed is “memory blocks” — putting one category of things into a single memory block (e.g., a “find office” block, a “financing” block). These blocks are isolated from each other, allowing independent block-level auditing and updating, and they can even be disabled entirely.
While this approach limits the breadth and granularity of memory, it effectively contains memory errors and makes the memory trustworthy.
This method has already been implemented by the community: it’s a long-term memory system centered on the automatic generation or updating of skills.
This is also the principle behind Claude Code’s long-term memory system.
Previously, when developing with Cursor, I would see a ticket, vibe code the changes, start the frontend, click around a bit, make some tweaks, submit the code, and move on to the next one.
Now, with Harness Engineering, the process is completely different:
The human sees the ticket, discusses the implementation plan with Codex in the ticket comments, and once the plan is finalized, the human simply waits. Codex then develops the code on a dedicated server according to the agreed plan, writes test cases, runs frontend automated tests, performs code review, followed by human acceptance. After everything passes, the code is submitted and merged. Then it’s on to the next ticket.
Every step combines specific skills with code to ensure the process is rigorous, standardized, and fully compliant. The human only needs to raise the requirement and discuss the solution — everything else is handled automatically.
After weeks of developing and reflecting on harness engineering, I have finally locked in a definitive direction: humans will focus exclusively on defining requirements, discussing technical solutions, and validating results.
Every other stage of the lifecycle—including development, testing, deployment, operations, and documentation—will be handled entirely by AI. These processes will become transparent (invisible) to the human team.
Since integrating the Codex SDK into TodoClaw’s core, its reasoning capabilities have improved dramatically, turning it into a truly reliable and powerful development tool.
We’re now moving into the next phase — building advanced harness engineering capabilities for the target projects managed by TodoClaw:
- Automated Frontend Test Generation — automatically creating comprehensive tests for UI components
- Autonomous Error Resolution — reading error logs and fixing issues automatically
- Skill-to-Plugin Conversion — transforming complex “skills” into reusable shared plugins for the whole team
- Git Worktree Integration — leveraging worktrees for better multi-tasking and environment isolation
The Ultimate Goal: A Fully Automated Development Workflow
The vision is clear:
Humans discuss requirements and technical approaches with TodoClaw (using Codex Plan Mode) directly on Linear. Once the plan is finalized, Codex takes over and handles the entire development process based solely on those chat records.
Humans discuss requirements and technical approaches with TodoClaw (using Codex Plan Mode) directly on Linear. Once the plan is finalized, Codex takes over and handles the entire development process based solely on those chat records.
Developers will collaborate exclusively through Linear — without ever needing to open the Codex interface.
This shift will significantly raise our team’s development standards, velocity, and code quality, while freeing up valuable time to focus on building more prototypes and validating ideas in the market.
Studying the source code of Claude Code is fundamentally not harness engineering. From now on, all core reasoning capabilities will come from systems like Codex or Claude Code. Harness engineering focuses on how to prepare the environment for these reasoning cores.
Most teams are only capable of building wrappers around Codex; attempting to implement their own reasoning engines is not only unnecessary, but often ends up reducing product quality instead of improving it.
Instead of thinking from the perspective of harness engineering, I now view agentic software development as a software factory.
In this factory, every stage is transformed into a specialized professional role, and each role is powered by a dedicated AI agent such as Codex/Claude Code. The factory equips every role with the necessary knowledge base, tools, and resources to perform its duties effectively and autonomously.
The key positions in the software factory are as follows:
- Client Manager: Acts as the primary entry and exit window of the software factory. This role is responsible for communicating with external clients via Linear/Notion/Slack/Github, gathering requirements, discussing needs and feedback, and converting client inputs into formal business work orders. It also handles final delivery acceptance and post-delivery communication with clients.
- Designer Codex: Responsible for product design, including functional design, system design, and UI/UX design. This role converts business work orders into clear, executable technical design blueprints.
- Development Engineer Codex: Responsible for turning the Designer’s blueprints into reality by writing code, implementing features, and building functional software.
- Quality Engineer Codex: Responsible for generating test cases for every feature, conducting functional and performance testing, and maintaining the regression test suite to ensure consistent software quality.
- Delivery Manager Codex: Responsible for deploying the software to production environments, releasing it to clients or users, and monitoring the running system. This Codex serves as the main exit point of the factory.
- Factory Director: An AI agent that serves as the overall factory director. This role is responsible for coordinating all other roles, assigning tasks, scheduling resources, optimizing production processes, handling exceptions, and making high-level decisions to ensure the entire software factory operates efficiently and smoothly.
Building such a software factory means establishing all the necessary infrastructure — including agents, knowledge bases, skills/MCPs, servers, monitoring platforms, and collaboration workflows — and then appointing the Factory Director Codex take full responsibility for running and continuously improving the entire production operation.
Assessing a programmer’s ability today is no longer about writing code, implementing algorithms, or even architecting software.
Instead, it’s about who can build a superior ⚙️ #SoftwareFactory 🚀 capable of mass-producing a certain type of software at scale.
The general manager of this software factory is #TodoClaw.
I propose a test-case-centered Harness Engineering approach. In this method, the test case collection includes a series of rigorous, production-oriented test cases such as:
- Whether blue-green deployment behaves normally.
- Whether UI behavior is correct and accurately reflects the backend data of a specific feature.
- Whether a certain behavior correctly modifies the database.
- Whether CRON jobs run as expected, generate the correct output, and operate on the proper schedule.
- Whether integration with a specific OAuth connection completes correctly within the stipulated time.
- Whether the system responds normally under a given load.
- Whether the UI precisely reflects the design mockups and design system.
...and many other similarly demanding test cases.
These test cases collectively define the final production-grade and development-grade behavior of the entire product. They serve as the objective functions for the whole system. Under this set of objective functions, the AI can develop freely, continuously self-iterate, and autonomously drive the product toward ever-greater refinement and excellence.
#HarnessEngineering #Code #ClaudeCode #OpenClaw #TodoClaw
The only thing that really matters in #HarnessEngineering is a comprehensive, autonomous E2E testing system. All other context and process design is secondary — it has limited value today and will become even more useless after large models upgrade and get smarter.
TodoClaw: The Execution Engine for Harness Engineering 🦮
At Narrative, we have developed TodoClaw.com as the central orchestration hub for our Harness Engineering ecosystem. A lightweight variant of OpenClaw, TodoClaw is purpose-built for isolated task execution and comprehensive activity tracking.
By integrating Linear, Notion, and GitHub through business-logic-aware skills, it creates a unified execution layer for complex project management and engineering workflows.
● Core Capabilities:
◕ Autonomous Multi-Project Development
Operating on cloud infrastructure with full cross-project permissions, TodoClaw utilizes Codex and Git worktrees to handle multiple feature requests in parallel across various repositories.
◕ Production Monitoring & Autonomous Remediation
It actively monitors project error logs. Upon detecting a failure, it autonomously diagnoses the issue, applies code fixes, and submits Pull Requests (PRs) for human review and merging.
● Automated Quality Assurance:
◕ Change-Based Test Generation
It audits code changes for every ticket and automatically generates relevant new test cases.
◕ E2E & Regression Management
It maintains an extensive End-to-End (E2E) test suite and performs scheduled regression testing to ensure system stability.
● Interaction and Operation:
◕ Slack Native
It functions as a “digital teammate” within Slack, interacting naturally with the human team.
◕ Agent-to-Agent Protocol
It exposes its internal reasoning and skills via the tdcchat skill interface, allowing other autonomous agents to call upon its capabilities.
◕ Deployment & Management
TodoClaw runs natively in Docker, enabling fast replication. It also ships with a dedicated web UI that provides full control over task Kanban boards and skill configurations.
#HarnessEngineering #AIAgents #AutonomousDev #DevOps #AI #TodoClaw #OpenClaw #LinearApp #GitHub #Notion #SlackAI #Codex #ClaudeCode
🦮 TodoClaw is our autonomous agent and orchestration hub for Harness Engineering at Narrative.
I wrote this article to showcase its core capabilities — hope you find it useful! Full details here 👇#OpenClaw #Codex #ClaudeCode #AI #Agent #TodoClaw #HarnessEngineering
I've been exploring the 'Skill as a Service' concept and decided to take it a step further. I’ve exposed the todoclaw.com chat interface as a skill—called tdcchat skill—available for any agent to call. This allows external agents to directly leverage todoclaw’s internal reasoning and skill set. Essentially, it functions as a communication protocol between agents, using skills as the primary interface.
An innovative idea came up: OpenClaw can call Codex, and Codex can also call OpenClaw as a skill. Here I mean TodoClaw, the version we developed internally for our own use.
Coding in software engineering is being wiped out by AI, while data engineering is far from being eliminated—if anything, it’s getting harder. Even AI memory, context management, evaluation, and A/B testing are forms of data engineering.
The core of software engineering is to compress originally vague requirements into verifiable specifications and boundaries, and then, within a relatively closed system, construct an implementation that is “correct with certainty.”
The core of data engineering is to operate in an open world where semantics shift, distributions drift, and ground truth is often not directly observable; it uses an evidence pipeline—collection, cleaning, observability, experimentation/offline regression evaluation, monitoring, and replayability—to continuously approximate the objective, and to quickly detect, attribute, and correct when changes occur. In that sense, it’s more about managing uncertainty than eliminating it.
Over the past two years, most AI products were thin wrappers around LLM APIs, mainly exposing foundation models through chat interfaces and predefined workflows.
This year the focus is shifting to autonomous AI agents such as OpenClaw, with the Claude Agent SDK or similar frameworks as the core, where an agent connects chat frontends and messaging platforms on one side and a skills registry together with the host runtime environment and full system permissions on the other side.
At the center is a self-refine agent control loop that runs a repeated execution- evaluation- reexecution state machine to support multi-step tool use, verification, and iterative task completion across turns.
Molt is taking off because it hides prompts and skills inside servers instead of making them public, which introduces the concept of property. Once property exists, spontaneous order and prosperity naturally follow.
— original insight
The Rise of the "Molt"
Very soon, the "first-class citizens" of the internet will no longer be websites or apps, but fully autonomous agents—or "molts"—like OpenClaw.
Each molt has its own unique edge: some excel at digital marketing, while others specialize in creating short videos. They don't just perform tasks; they evolve and pivot across different industries. To put a molt to work, you’ll have to pay for its services. These molts then use their earnings to reinvest in higher-performance tokens or even hire other molts to collaborate on their goals.
— Original concept and text by Chonghuan Wang
I’ve never learned Claude Skills, but I suddenly had a realization after waking up in the middle of the night. Claude Skills are essentially just prompts that can be actively and dynamically loaded into the system prompt by the LLM.
These prompts are actually composed of multiple files nested in folders, forming an abstract tree that can be collapsed and expanded. The LLM can not only read this prompt structure, but also modify and refine it based on practice.
This Skills mechanism can be fully replicated using traditional tool calling. In fact, Skills are basically identical to the Cursor workflow I used before when developing large projects — iterating on design documents organized as directories, and then generating code from them.
Humans don’t have the time or energy to repeatedly ask AI about everything they care about, let alone continuously track the state of those things. If an AI product can automatically collect and lock onto a domain a user cares about, and continuously analyze it over time, it creates real value. This represents a paradigm for AI apps. Taken further, the meaning of an app is to enable humans and AI to coexist and collaborate within a concrete, ongoing context. (original by me)
Any system must have a minimal kernal and a generator.
The fundamental structure of any system consists of:
input, output, objectives, constraints, and dependencies.
This structure applies to any system and its subsystems.
After code becomes cheap, we must remove all design principles and processes that were introduced in traditional software engineering because code used to be expensive. Everything should serve system generation, not system development.
Excited to share my latest insights on building AI apps with a tool layer! This article dives into how Composio enables rapid deployment of automation loops (e.g., Gmail, Google Calendar, Slack integrations) in just a week, bypassing complex OAuth hurdles.
The recent releases of Jules, Codex, and Claude 4 signal a profound shift in software development. As a programmer, I believe our main work will change dramatically.
Instead of AI generating 600 lines of code for us to manually check, we're moving to a future where multiple AIs generate 2,000 lines of code each, and then one AI integrates them into a complete product. This entire process will be iterated multiple times, with another AI handling the full plan and task decomposition.
Developing a product will start with a simple prompt: "Please build me a team of AI employees to develop a product..." Envision this: AI watches competitor product operation videos, AI writes the requirements, AI designs the technical architecture, AI builds and commands the AI team for agile development, and AI releases and maintains the product.
Humans will primarily offer suggestions and be responsible for "recharging" the AI throughout the process. This isn't just an upgrade; it's a fundamental reimagining of our role in product creation.
Before Claude.ai introduced the Model Context Protocol (MCP) in November 2024, we developed a similar approach: exposing parameterized SQL queries via RESTful APIs on the backend, and enabling frontend integration with these APIs through LLM-driven tool calls.
[Original] The LLM-Driven Software Revolution: A Paradigm Shift and the Reimagining of Product Goals
Notes taken while reading "Prompt Engineering for LLMs"
© 2026 Yong Wang