The Helm Desk Suitelet inside NetSuite: a morning summary showing 28 agents ran 28 times for $5.73 with 45 findings needing a decision, a critical Sales Order Health finding open with the agent's evidence on the left and a live NetSuite evidence panel on the right, and a row of decision buttons below

Helm is a framework on which AI agents can run inside NetSuite. My test instance has twenty-eight of them at the moment. Each one wakes at a set hour, reads a slice of the ledger under a role that I assigned, and logs what looks wrong.

The framework handles the things that make unattended agents potentially dangerous. It enforces a policy on every tool call, limits the number of turns they can make, as well as the tokens used, the cost and writes per run. It halts the run when a cap is hit, screens tool results for prompt injection, records a full transcript, and deduplicates findings so the same problem isn't raised every morning. The agents concentrate on their work, while Helm concentrates on keeping them inside the lines.

That reliability is exactly what I set out to build.

What I didn't set out to build, and have come to think is just as important, is that Helm is an ordinary NetSuite application. It's really just SuiteScript. Its agents, runs, findings, drafts and policies are custom records. There's nothing about it that NetSuite doesn't already know how to search, report on, secure, or extend.

Embedded, so nothing is hidden

Helm could have been an external service. Most agent products are: something outside the account pulls data through an integration, reasons about it elsewhere, and pushes results back. Every one of those seams is a place for data to be stale, for permissions to drift, and for "what did the agent see?" to become a question about someone else's logs.

Inside the account there is nothing to sync. An agent sees exactly what its role sees, at the moment it runs. The only thing that leaves NetSuite is the conversation with the model, and every turn of that conversation comes back and is stored on the run record along with the records the agent touched, the tokens it spent, and what it cost. When a reviewer asks why a finding was raised, the answer is on a record they can open.

The same is true of the instructions. Each agent's system prompt is a field on the agent record, with a version history. If Ledger Sentinel flags a duplicate vendor payment, then you can open its prompt and read the exact instruction that produced the flag. If you think that the instruction is wrong, you can change it, and Helm will tell you what changed and when. I think that this matters more for agents than for most software. An agent makes judgment calls on your ledger at 3 a.m. If all you can inspect is its output, then trusting it is an act of faith. If you can read its instructions and its transcript, then trusting it becomes an ordinary engineering decision.

Ordinary records, so anyone can build on it

Here's where the "just another app" quality started paying off in ways I hadn't planned.

Helm's findings are custom records. A saved search reads them. A portlet displays them. A workbook joins them. A Suitelet can pull in the transaction each finding is about. So when I noticed a problem with how findings were being reviewed, the fix didn't need a new API or a change to any agent.

The problem was this. I'd built one console for running the fleet - prompt versions, policies, budgets, halt reasons - and that console was also the only place a finding could be reviewed. It's the right tool for me. But it's the wrong tool for the person who owns payables.

So, using Sonar, I built a second Suitelet, which I'm calling Helm Desk. It reads the owner field that already exists on every agent record and shows the logged-in user only the findings from agents they own. The AP lead sees payables findings. The warehouse manager sees inventory findings. One deployment, and the scope comes from the data rather than from per-user code.

What makes the Desk worth having is the evidence panel. A finding about a duplicate payment used to be a paragraph of text with record IDs in it. In the Desk, that finding renders both payments side by side, pulled live from the transaction tables, with the bills each payment was applied to and a fresh check for any other same-amount payment to that vendor in the past 30 days. When I opened the first flagged pair, that live check also surfaced a third matching payment from earlier in the summer, which the agent's brief hadn't asked it to look for. The agent did what it was told. The Desk, sitting on the same tables, could ask a slightly wider question, and the two together gave the reviewer more than either would alone.

The agents are one consumer of Helm's records. The Desk is another. A controller who wants a weekly email of unreviewed critical findings can write a saved search and schedule it. An admin who wants to know whether the agents are paying for themselves can join run cost to findings in one query. None of that touches an agent, a prompt, or a policy, and none of it needed me.

The one rule that openness demands

If you can read the tables, then you can write to them, and some writes have consequences that aren't obvious from the schema. Helm's deduplication keys on the reviewed flag. The morning brief reads disposition tallies. A dashboard that stamped those fields directly would break both, and it silently.

So the Desk records decisions by calling the same server-side function the operator console uses, and I've written that into the design document as a rule: read the tables freely, write through the code that understands them.

I don't think that rule is specific to Helm. Anyone building on a system they didn't write, inside NetSuite or anywhere else, has to make the same distinction.

Wrapping up

If you're building agents that act on a system of record, then put the agents' own records in that system, in the plainest form it supports, and make the instructions readable by the people whose work the agents examine.

You might be giving up some elegance. But what you'll get back is the ability for anyone with access to look at what the agent did, question why, and build something new on top of it without needing to ask you first.