From Bug Hunting to Engineering Methodology: Lessons from AI-Driven Development

In a previous blog, I shared what I learned from building an internal C++ library used by the Arteris verification team to generate constrained test stimuli. To summarize, AI is an information-retrieval tool; verification stays with the engineer; and a large portion of the information in our job is much easier to verify than to find.

I first put this thinking into practice with bug hunting. While developing the library, I used AI to search for scenarios that would trigger failures I could objectively verify. That worked well enough that I began applying the same principle more broadly, first to other types of bugs and then to the library design. What began as a way to find software defects ultimately changed how I thought about AI-driven development, not just for testing, but for software design and engineering research, too.

Use AI as an adversarial sparring partner.

My first use case was deliberately simple. I leveraged my experience to identify incorrect behavior, then asked an AI to find a scenario that would trigger it. There’s no guarantee it will find one, but when it does, I can verify the result rather than debate whether the model is right.

Take segmentation faults. As the author of the library, I know it must never segfault under valid operating conditions. This gives me a wonderfully simple criterion: “The program segfaults, yes or no.” From there, I can build a test harness around that criterion and let an LLM search for inputs that pass or fail the filter.

At the end of one long segfault hunt, I had tens of tests that triggered segmentation faults. Some cases overlapped, but every one was valid and worth investigating. Then, before I’d even had a chance to fix them, a colleague sent me an example where the library segfaulted. A week later, after I’d addressed the issues the AI-driven tests had uncovered, I tried his example again. It no longer crashed, even though I had never investigated that particular case.

I then applied the same approach to memory leaks and found several bugs. At that point, it was tempting to summarize the method as: “LLMs as an adversarial sparring partner.” And that is part of it, but the more interesting lesson came next.

Design the software to make failure verifiable.

The Arteris C++ verification library above generates legal stimuli from a data model and constraints between fields. Generating a valid sample can be difficult because the software must satisfy many relationships simultaneously. Checking the result is much easier because once a sample exists, I can automatically verify whether it respects the constraints.

That asymmetry changed how I thought about the architecture. Rather than only checking outputs in an external test environment, I made verification part of the library’s default behavior. After each sample generation, the library runs a verification pass; if the sample is not legal, it raises a defined exception.

Now I had another objective criterion for AI exploration — to find a scenario that triggers these same exceptions. The resulting hunt uncovered a small number of additional bugs, exercised from different angles by a dozen tests. Again, the important point was not that the LLM understood whether the verification library was correct. The software itself defined correctness, AI searched for exceptions, and I investigated the confirmed failures.

This is where AI-driven development begins to influence software architecture rather than simply testing. If a behavior can be made objectively verifiable, build that verification into the system and give AI room to search for the conditions that violate it.

Verification cost tells you where AI can help.

From there, I started looking at problems where verification takes a little more work. For instance, I could look for performance bugs by finding cases where sample generation took far too long for a relatively simple model. I could also search for potential problems caused by undefined behavior, inconsistencies in the documentation, or issues in the build flow. Still, none of these is quite as simple as “the program segfaults, yes or no?” However, the same principle still applies: they are much easier to verify than they are to find.

These cases are different, but I think about them on the same continuum. The easier a result is to verify, the more freedom I can give AI to explore. As verification becomes more expensive, subjective, or dependent on broader system knowledge, I need to spend more engineering effort evaluating the results.

This has changed the way I think about where AI can help. What matters is not simply whether AI can tackle a particular problem, but how easily I can verify what it finds. Prompting matters, but having a reliable way to check the result often matters more.

Design research is still information retrieval.

So far, this sounds like a testing methodology. The above library itself was also designed with extensive help from AI, which pushed the idea further. Design research is still information retrieval.

Think about how much of an engineer’s day involves searching. We investigate APIs, compare algorithms, read specifications, explore architectural alternatives, look for implementation strategies, and learn unfamiliar technologies. An LLM can accelerate that work because it does more than retrieve an existing document. It can combine information and propose an implementation or alternative that may never have been written down in exactly that form.

That does not make the proposal correct, nor does it transfer the design decision to the model, but it does change the amount of time required to reach something worth evaluating. Instead of spending hours identifying possible approaches, I can spend more time comparing them, testing assumptions, understanding trade-offs, and deciding what belongs in the design.

For me, that is a much more useful framing of AI than treating it as an engineer who produces answers. It is a formidable information-retrieval tool, and retrieval is a very large part of R&D.

Measure the research-to-verification ratio.

One practical exercise I would suggest to other engineers is to measure the research-to-verification ratio of your efforts. How much time do you spend finding an answer or a scenario, and how much time do you spend determining whether it is correct? Wherever the search dominates and verification is comparatively cheap, there is a strong opportunity to use AI to accelerate the work.

I would also classify problems by the level of required reliability. Some engineering outputs demand extremely high confidence. Others only need to be credible enough to justify the next experiment or point you toward a promising approach. Treating every AI interaction as though it must deliver a definitive answer misses that distinction.

This doesn’t lower the standard of engineering or hand responsibility over to AI. Rather, it changes where I apply my engineering judgment. AI can accelerate the search, but I still have to verify what it finds, understand the trade-offs, and decide what is useful.

Change where engineering effort goes

As AI models improve, they will be able to find more information, tackle more complex problems, and produce increasingly reliable results, expanding the range of engineering work they can help with. But the underlying principle remains the same: engineers still need to verify what AI finds and decide how those results should be used.

At Arteris, this experience is helping shape a practical approach to AI-driven engineering: build verification into the development process, give AI well-defined problems to explore, and keep engineering judgment at the center of the decisions that follow. The goal is not simply to add AI to existing workflows, but to structure those workflows so that AI can accelerate the search for answers without compromising the rigor needed to validate them.

That is the methodology I took away from building the above library. Use AI aggressively where it can search faster than you can, but give it objective filters whenever possible. Architect software so that important behaviors can verify themselves and reserve engineering expertise for the work that requires it most, such as interpreting evidence, understanding trade-offs, and making decisions.

A huge amount of the information in our job is much easier to verify than to find. Once you start looking at engineering work through that lens, AI becomes less mysterious and more of another tool for changing where we invest our time.


Explore Arteris IP:


×
Semiconductor IP