SimplifyC++ Article
When Performance Looks Promising but Trust Falls Short: My Experience Using AI to Build MyLang

I built MyLang on top of MyLangAssembler, aiming to create a systems language that combines power and efficiency with a clear, dependable design. I had no team of developers to share the work on its major components, so I used AI extensively. It helped me design and implement many parts of the language, and the pace of progress was remarkable. I then tested programs written in MyLang and compared their performance with equivalent programs in C, C++, and Rust. In the tests I ran, the results were close, sometimes slightly faster and sometimes slightly slower, depending on the case.
That was encouraging, but it taught me a lesson more important than the performance figures: a compiler can produce fast programs in familiar tests without being a compiler I can trust or release to the public.
Where Did Confidence Break Down?
A compiler's job is not merely to produce an executable that works for ordinary examples. It must preserve the program's meaning through every stage: parsing and type checking, intermediate representation and optimization, instruction generation and linking, memory handling, and the system interface. A bug in a rare case may remain hidden until a particular execution path meets a different optimization level or a memory-use pattern absent from the initial tests.
In the previous implementation of MyLang, I came to realize that critical components needed a documented rebuild and full human review, particularly the IR, Shared IR, memory management, and the boundaries of undefined behavior (UB). These are not details whose validation can wait until the end of the project. If their rules are not explicit, and the transformations between them lack sufficient justification and testing, the language may give convincing results today and fail tomorrow in a program someone has trusted it to build.
For example, it is not enough for the translation of a conditional or loop into IR to look reasonable when reading the code. It must preserve the correct values along every execution path. Nor is it enough for a compiler to optimize memory access successfully in a small test: the optimization must rest on explicit rules about data lifetimes, aliasing, and the assumptions the optimizer is allowed to make. A performance test measures what happened in particular runs; it does not establish that every transformation is correct or every possible program is safe.
So I made a difficult decision: I removed the previous implementation of the language and began building MyLang again myself. That does not mean everything produced before was worthless, or that the benchmark results meant nothing. It means I could not put my name on a language whose fundamental decisions I could not trace and defend, stage by stage.
What Can AI Do Well?
My experience does not diminish AI's capabilities. It has been a powerful assistant for exploring alternatives, explaining ideas, drafting initial implementations, converting large amounts of data and code, suggesting test cases, and speeding up repetitive work. For a large project carried out by one person, that help has real value.
But there is a difference between producing code that appears correct and building a chain of evidence that can be examined. AI may write components that are locally coherent while their assumptions conflict across the system. It may also propose a test that confirms the same behavior it assumed in the implementation, creating reassurance without checking an independent specification. The more compiler stages depend on one another, the more important it becomes to trace decisions, explain why each transformation is sound, and know when to reject a program instead of generating an output with uncertain behavior.
As the project's owner, that responsibility is mine. The tool does not decide whether to publish a compiler or declare it ready. I define the specifications, review the implementation, and decide whether the evidence is sufficient. In the new MyLang, I therefore use AI as an assistant and reference, with its contributions tied to tests and thorough human review. I design the foundation and make the decisions on which the language's safety depends.
Why Is MyLangAssembler Different?
It is important to distinguish between the language and the assembler within MyLang Toolchain. I designed most of the foundation of MyLangAssembler by hand. AI provides substantial help with research, converting code, data, and instruction information, and accelerating work at scale. I then examine what I need and incorporate it into the design myself. My direct involvement in the assembler's structure and fundamental decisions is different from my relationship with the earlier language implementation.
That distinction does not exempt the assembler from scrutiny. Correct instruction encoding, symbols, branches, linking, and edge cases require ongoing tests, comparisons, and review before any release can be described as ready for the public. It does explain why I continued developing MyLangAssembler while choosing to dismantle and rebuild the language implementation.
How Will I Move Forward?
I will start with written specifications that define the meaning and limits of each feature, then implement them in small, reviewable stages. Each stage will have tests for correct results, tests for inputs that must be rejected, edge cases, and independent comparisons where comparison is appropriate. I will treat memory rules, undefined behavior, and intermediate transformations as part of the design itself, not additions to make after the compiler is complete. When a fundamental decision changes, the documentation and tests must change with it before I claim that the feature is stable.
This experience does not show that AI is unsuitable for systems tools, or that speed and benchmark results have no value. It shows that faster code generation does not shorten the work of establishing correctness. For tools people will use to build their own software, one of the most important skills may be recognizing where a working example is insufficient, and being willing to stop and rebuild when the grounds for trust are not there.
I will continue using AI because it is an immensely useful assistant. But designing the foundation, understanding its assumptions, independently validating it, and taking responsibility for its release will remain human work in MyLang Toolchain. I want to build a language and assembler others can rely on, rather than an impressive performance demonstration hiding questions I have not yet answered.
What did you think?
Sign in to react or comment.
Comments
0No comments yet.