I’m a software engineer, and I built the first version of Scorebook myself.

The idea came from years of volunteering at the volleyball score table. I knew the work, understood the frustrations, and had the technical background to start building something that could help.

As the app grew, so did the work ahead of me. Multiple people needed to work on the same match. Their changes needed to survive connection problems. The interface needed to work comfortably on phones. Every improvement needed testing, and I needed a dependable way to get changes into users’ hands.

My estimate is that working through all of that on my own would have taken another year. With Claude Code and Codex, we covered that ground in less than three months.

That still amazes me, even knowing how much time I’ve put into leading the work.

Today, Scorebook has a small development team: me, Claude Code, and Codex. I’m the lead designer and product owner, and I remain involved as a software engineer. I bring the volleyball experience, evaluate technical proposals, work through the user experience, and decide what reaches users.

Claude and Codex investigate problems, propose solutions, write code, build tests, and review each other’s work. They’ve helped me take an app I had already started and develop it much further, much faster.

Together, we’ve also built the process around it: a GitHub ticketing system, automated testing, independent reviews, and separate environments for testing and releasing changes.

Keeping that team productive takes active involvement. I need to give the agents enough room to solve problems while recognizing when a technically reasonable solution is taking the product in the wrong direction.

Giving everyone a job

I don’t need to tell Claude or Codex how to write every function. I do need to make the intended outcome clear.

A request might start with something I noticed on my phone: the number picker covers the field, a substitution takes too many taps, or a screen makes me stop and figure out what to do next.

From there, an agent can inspect the existing application, trace the cause, suggest an approach, and implement it. It can also write the tests and prepare the change for review.

The other agent then examines the actual changes. If Claude builds it, Codex reviews it. If Codex builds it, Claude reviews it.

That second opinion has been useful. During one proposed scoring redesign, reviews uncovered problems involving two devices claiming the same scoring role, actions getting lost when entered quickly, and pending work that wouldn’t survive a refresh. Those were meaningful problems in how the design would behave.

The agents can challenge each other’s reasoning and work through corrections. I still decide whether the result solves the right problem and whether we should release it.

We keep the work on a board

We use a GitHub project board with a straightforward flow: Todo, In Progress, In Review, Done.

Each ticket describes a specific problem and what a successful result should look like. That gives the agent something more durable than a request buried in a long conversation.

When an agent starts a ticket, it works in its own isolated copy of the code. That matters when both are working at once. One can improve a screen while the other investigates an unrelated issue without overwriting unfinished work.

When the change is ready, the agent opens a pull request. That packages the proposed changes, the reason for them, the testing results, and anything the reviewer needs to examine closely.

The ticket stays in review while we work through approval and release. Code being written is only one part of completing the job. I want to know that the change reached the intended environment and behaved correctly there.

We also keep the rules, architecture, and working agreements in the repository. The agents read those documents when they begin work. When we make an important decision, recording it gives the next session something concrete to follow.

Otherwise, I would spend too much time explaining decisions we had already made.

We built a release process around the team

The overall process is our software development life cycle: decide what matters, describe the work, build it, test it, review it, release it, and learn from using it.

The automated part is often called CI/CD. For Scorebook, that means proposed changes go through a repeatable set of checks before they can move toward release.

The system builds the application, runs the tests, and exercises workflows in a browser. Some checks simulate several people working on the same match. Others test saving, permissions, or what happens when a device loses its connection.

After review and my approval, a change moves to staging, our separate environment for checking it before it reaches users. Production requires another approval, and we check the deployed application afterward.

Claude and Codex do much of the work involved in preparing and verifying those steps. They don’t get standing permission to release whatever they finish.

That distinction matters with agents that can operate development tools directly. Writing a change, modifying a database, and deploying to production have very different consequences. Our instructions make those boundaries explicit.

We improved the process when it started slowing us down

As our testing grew, waiting for results became a bigger part of development.

We looked at what the automated process was doing and found opportunities to reorganize it. Independent checks could run at the same time. Browser tests could share one prepared build instead of each rebuilding the application. The larger group of tests could be split across two separate test environments.

We even tried dividing that work three ways. It didn’t shorten the overall wait because another part of the process still took longer. We kept two groups rather than paying for extra work that didn’t get us an answer sooner.

We also removed unnecessary repetition between reviewing a change and deploying it. The deployment checks still verify the exact version being released and confirm that it came through a successfully tested pull request. They don’t repeat every expensive browser scenario solely because the approved code moved to the next stage.

Those improvements matter when you’re making frequent adjustments. I can try a change, give feedback, and get to the next version sooner.

The agents helped build and improve that entire system. Their contribution extends well beyond the screens people see.

Knowing the job changes the decisions

There’s a point in product development where experience helps you recognize a problem before you can fully explain it.

I can try an interaction and know that it asks too much of someone keeping the book. Then I have to turn that reaction into useful direction: this control needs to stay visible, that step is unnecessary, or this information belongs next to the action.

The number picker was a good example. Showing available roster numbers sounded convenient. On a phone, the extra content made the picker larger and interfered with the field being edited. We removed that list. Later, we brought back a much smaller shortcut for a specific returning player.

We had to work through which convenience actually helped during use.

My software engineering background matters here too. I can examine how a proposed change fits into the application, question its assumptions, and understand the consequences of replacing something that already works.

That gives me two ways to evaluate a proposal. I can look at it as an engineer and then use it as a bookkeeper. Sometimes I agree with the technical reasoning and still decide the result asks too much of the person at the table.

A larger example involved redesigning how scoring actions were saved. We explored a more elaborate approach that offered stronger control over individual actions and which device could submit them.

There were legitimate reasons to consider it. But the transition introduced more work and more ways for things to go wrong. An early version even exposed a control for adopting the new scoring method. A volunteer should never need to understand an internal change like that just to keep score.

We eventually backed the experimental foundation out of staging and continued with the existing scoring approach. The experiment never reached production.

That’s a decision I need to remain involved in. The agents can explain the architectural benefits and investigate the risks. I have to weigh those against what the product needs now, what is already working, and what the change asks of its users.

Having an engineer and an experienced bookkeeper leading the work helps me judge where to trust their approach, where to ask for evidence, and where we need to change direction.

The productivity is remarkable, but it needs direction

There are parts of this arrangement I find tremendously valuable.

I can ask for another design iteration without worrying that the agent is tired of revisiting the same screen. I can ask it to investigate an awkward failure, document the cause, add a test, and run the checks again. Repetitive work doesn’t need to wait until someone feels like doing it.

That makes thoroughness much more affordable.

It doesn’t make thoroughness automatic. An agent can write a test that confirms its own mistaken assumption. It can fix the example I described while missing another place where the same problem occurs. It can produce a convincing explanation before it has gathered enough evidence.

We’ve learned to look at what the test actually proves. Does it reproduce the failure? Does it use the real saving process when saving is the concern? Does it check what the user sees, or just confirm that a page opened?

We’ve also had to be careful with parallel work. Two agents changing the same files can create extra coordination instead of extra progress. Separate workspaces help, but sometimes I need to sequence the work.

And there are ordinary service limitations. Usage limits, interrupted sessions, and expired logins can stop progress. We recently had an independent review pause because the reviewing agent’s login needed to be renewed. Having the workflow written down meant we could see exactly what remained unfinished.

The agents don’t get tired in the human sense. They can still lose context, miss a requirement, or pursue an unhelpful approach for a long time. I need to watch the result, not assume persistence means progress.

The cost changes what I can take on

Hiring two software engineers would be a substantial commitment.

For a rough U.S. benchmark, the Bureau of Labor Statistics reports a median software developer wage of $135,980 in May 2025. Two salaries at that level would total about $272,000 a year before benefits and other employment costs. Actual hiring costs would depend on experience, location, and the roles involved. BLS software developer wage data

My access to these two AI engineering tools comes through monthly subscriptions. That is a very different financial commitment.

The subscriptions aren’t the entire cost of Scorebook. There’s hosting, supporting services, and a considerable amount of my own time. I’m supplying engineering judgment, product leadership, design direction, volleyball knowledge, testing, and release decisions. The services also have usage limits.

I wouldn’t describe two subscriptions as equivalent to employing two experienced people with independent responsibility for the product. But the engineering work I can get done with this arrangement has changed the scale of what I can build.

I had already invested my own engineering work in getting Scorebook started. The subscriptions gave me the capacity to take it further without hiring a development team.

My comparison is based on the work I knew was still ahead of me. What I expected to spend another year building and working through took less than three months with the agents. My time and experience were still essential, but I could spend more of that time directing, evaluating, and refining the product while implementation, testing, and review moved forward.

For me, that combination of cost, speed, and capability is extraordinary.

I still need to know what I want Scorebook to do, recognize when something feels wrong, and be willing to stop or redirect the work. When I do that well, Claude and Codex give me the ability to follow through on ideas at a pace I couldn’t sustain alone.