EnriqueMark
Design Patterns & Architecture

Two sides of testing

2026-09 English

As far as testing itself goes, deep coverage is fine. But it isn’t enough to keep a program running solidly. Tests might cover the function body’s own logic in detail, the output it produces, even the whole e2e. But the input itself is always built on assumptions made up front.

Take assumptions about input data. I talked about this before when I wrote about property-based testing: input has a blind spot. And here, the best fit is to bring in PBT.

Classic test setups attack the thing under test with fake data, or with data built by people (these days, by AI). The obvious problem here is “how do I test the parts I can’t think of?” Even if you hand it to AI, it’s still affected by its context and still leaves things out. Even with the unit tests and e2e for the function body done in detail, an unexpected bug still showed up in the end. I stepped in this hole. After the post-mortem, I found the reason it happened was “dimension”.

I found my tests were focused too much on “logic and output” and ignored “input”. That caused the problem above. The logic was fine, and e2e with regular data under expected conditions was fine too, but when it really ran it hit a case that was “very common” and that I hadn’t tested. However deep you dig in a single dimension, what isn’t covered is still not covered. The key to this problem is “data”. The input data was a case I didn’t expect when testing. Part of that “didn’t expect” is of course the limits of my experience, but structurally it’s the same thing as “not knowing what you don’t know”.

So I went looking for a fix, and that led to PBT. I’d never formally brought it in before, because it didn’t seem that necessary and it’s a hassle, you have to add another package. Could I bring in a lightweight version, taking only the idea and leaving out the heavy setup for now? With that in mind, I started shoring up the input side. Even so, how to build the generator is something to think through. If the data is wrong, no amount of hitting will test anything real. At first I thought of building the distribution from existing data, but then it hit me that this is useless for a new feature. What do you do when there’s no data? So I can’t take concrete data. I have to take the structure: on top of the structure, what could it produce? And using PBT here has another problem. The values here may not be continuous values. They may come from different states crossing each other. How do you pin down which states to test?

At this point I have to bring in another testing concept, combinatorial testing, which is made for exactly this kind of scenario with combinations of multiple input parameters. The first thing combinatorial testing asks is that the tester clearly define what each state means. In other words, pin down the current state’s boundary, or how big its space is. “Define” here doesn’t mean the type. It means clearly defining its “business logic”, and what that state is actually for in the logic. Say there’s a state A. I need to spell out what state A means concretely, how state A should be handled under this input, and why it should be handled that way. Just going through this process removes a lot of the potential space.

Once the states are defined, you start building test data around the combinations. Classic combinatorial testing uses pairwise combinations (with n = 2 this is the full set of combinations). As the article linked above says, it’s the most cost-effective method, especially for preventing combinatorial explosion. When there aren’t many states, stopping here is enough, and a lot of problems caused by state combinations get caught. But then there’s a new problem. What if the values a state produces can’t be enumerated? Things like time or money. This is where PBT comes in, and it’s also where I brought PBT in in the lightest way.

Build a generator directly out of the combinations from before. But then another new problem comes up. It’s hard for me to assert from concrete values what counts as correct. I can’t “predict” the result, and PBT needs a gold standard to judge results by. So here I have to use metamorphic testing. When you can’t predict the output, you assert relations, and judge whether the output is valid by whether the results satisfy those relations. And this happens to be one of the most common kinds of assertion in standard PBT. With the prep done, you can start hitting it at scale. At the end you catch the results and use shrinking to narrow them down.

Once the tests are built, of course you run regression. I added this new input-side reinforcement to the regular tests, and specifically had an agent with no context run the whole suite, to see if the “routine tests” that run after every round of development could reproduce the problem I’d found before. (By the way, while building it, the agent baked the thing under test into the tests and wrote them as a regression. That’s classic shooting the arrow first and painting the target after. It can’t catch the real result, so HITL is needed to give clear feedback and oversight.) And it caught it. Looking back at the results afterward, this setup does work.

I think the most useful thing this problem exposed is about “process”. Like I said before in process-based quality control, my personal philosophy for using AI has always put process first. Modern manufacturing has shown countless times how much process contributes to finding and fixing errors. So compared to one local problem, I care more about “why didn’t my process find it”, as an early warning mechanism. Once a problem shows up, besides building a regression for the individual bug, going back over the process is also a necessary part. A bug fix can be about now, but the value of regression is in the future. A good quality control process is the same. It’s about now, and it’s about the future too.


Translation note. I wrote this in Chinese. This English version is an LLM translation, so the wording is not mine even though the thinking is. Original: 测试的两面.