The audit — EU AI Act compliance, AI governance, AI automation

An agent you cannot explain is an agent you will eventually switch off.

Novus Point builds custom AI agents for B2B operations, from London and Dubai.

We build them to be documented, logged, overseen and reversible — because an automation has to survive three separate encounters, and none of them is the demo. It has to survive a regulator asking what it does. An auditor asking to see the record. And month three in production, when the novelty has worn off and the thing is simply expected to work.

Request a 30-minute diagnostic What we build

Every engagement starts with the audit: a 30-minute diagnostic at no charge, then a written position, then engagement if it is warranted. Three areas sit under it — EU AI Act compliance, AI governance, AI automation. Senior-led throughout, by a firm that runs its own AI systems in production.

AI automation — one of three practices

The credibility base

We build this for ourselves first.

Novus Point does not only advise on AI. The firm builds and runs its own deployments in production — Hadar AI, a CRM for Dubai real-estate brokerages, and ADOZ.

That matters for one reason. Every position we hold on logging, oversight, evaluation and failure handling is a position we have had to live with on our own systems, at our own cost, in front of our own users. Advice that has never been operated is a theory. It is usually a tidy one, and it usually breaks the first time someone asks who is accountable when the model returns something confident and wrong.

Who leads the work

Jakub Piórkowski

Founder & Principal · London & Dubai

The firm is led by its Founder & Principal, Jakub Piórkowski. He is Chairman of the Supervisory Board of Carlson Investments SE, listed on the Warsaw Stock Exchange (WSE: CAI), elected in August 2026, and the founder of Buildeo Limited, which has spent eleven years in the UK market — construction and fit-out consulting, a different discipline, but the same habit of working inside regimes where the paperwork has to match what was actually done. His executive education in this field is short-course, not academic: AI-Driven Leadership at Stanford (2025), the Oxford Artificial Intelligence Programme at the University of Oxford, Strategic Thinking for the CXO at the University of Cambridge (2024), and Venture Capital at Imperial Business School (2023).

Engagements are led at principal level and stay there. Where the work needs legal, security or data expertise, the firm brings in advisers who work under our direction and on our responsibility.

We publish no client work, no metrics and no testimonials. We cite only what the Regulation says and what we have run ourselves; everything else on this page is our judgement, offered as such.

What we build

Four kinds of agent. Each one anchored to a business function, not to a technology.

The word “agent” has been stretched until it means very little. In practice, almost every useful build falls into one of four shapes. Which shape you need is a question about your operation, not about the model.

01

Process automation agents

An agent that carries a repeatable operational process from trigger to outcome: intake, validation, the sequence of system calls, the exception path, the handover back to a person. The value is not that a model is involved. The value is that the process stops depending on whether a particular person remembered a particular step. These are the builds with the clearest payback and the strictest need for an audit trail, because the agent is now acting inside a workflow someone signs off on.

02

Document and knowledge agents

An agent that reads what your organisation already holds — contracts, specifications, policies, correspondence, technical files — and answers questions against it with a citation back to the source. The engineering discipline here is refusal: the agent must be able to say that the answer is not in the corpus, rather than produce a plausible one. An unsourced answer from a document agent is worse than no answer, because it will be believed.

03

Client-facing assistants

An agent that speaks to your customers, applicants or counterparties. This is the highest-exposure category, and the one most often built with the least care. It carries transparency duties, a reputational surface, and a hard boundary question: which statements is this agent permitted to make on your behalf, and what happens when someone tries to talk it past that line. We design the escalation route to a human before we design the conversation.

04

Decision-support agents

An agent that assembles evidence, scores options and puts a recommendation in front of a person who decides. The design constraint is that it must stay support. That means the recommendation arrives with its reasoning and its inputs attached, the human reviewer is genuinely able to disagree, and the record shows what the agent proposed and what the person actually did. Automation bias is a known failure mode. It is designed against, not hoped away.

Governance by design

Governance is an architecture decision. It is not a document you write afterwards.

Most builders bolt governance on when someone senior finally asks a question. By then the agent has no logs worth reading, no defined oversight point, and no way to reconstruct why it did what it did on a Tuesday in March. Retrofitting that is expensive. Sometimes it is not possible at all, and the system is quietly retired.

Building an agent that can be documented, audited, explained, overridden and switched off means seven concrete things. We do them at the start, because they are cheap at the start.

  1. 01

    A written system description before the first line of code

    What the agent is for, what it may touch, what data it sees, where that data rests, which decisions are inside its authority and which are not. One document, in plain language, that a general counsel can read without a translator. If this cannot be written clearly, the scope is wrong and no amount of engineering will fix it.

  2. 02

    Logging that answers questions

    For every consequential action: the input, the model and prompt version in force, the tools and data sources called, the output returned, who saw it, and what they did next. Retention is set deliberately and in writing — a decision, not a default.

    There is a statutory floor coming, and it is worth designing to now. From 2 December 2027, deployers of the high-risk systems listed in Annex III will have to keep the logs automatically generated by the system, to the extent those logs are under their control, for a period appropriate to its intended purpose and of at least six months (Article 26(6)). For high-risk AI embedded in the regulated products covered by Annex I, the same regime applies from 2 August 2028. Neither date binds anyone before then. Both are impossible to satisfy retrospectively, which is why the log schema is a design decision rather than a compliance task.

  3. 03

    Named human-in-the-loop points

    Which decisions the agent may take alone, and which it may only propose. Who holds the override. Whether that person has the competence, the training and the authority to use it — which is the standard the Regulation will set for human oversight of high-risk systems, and a reasonable standard to hold yourself to whatever category you fall into. Oversight assigned to someone with no power to stop the process is not oversight.

  4. 04

    An off switch that has been tested

    A stop that brings the system to a halt in a safe state, plus the answer to the question nobody asks until the day it matters: what does the business do for the next four hours without the agent. The fallback is part of the build. We test the stop before go-live, not during the incident.

  5. 05

    Evaluation before production, and evaluation in production

    Before: a fixed test set with known-correct answers, including the adversarial cases and the edge cases the process actually throws up, run against every change. After: continuous sampling of live behaviour against the same standard. Model behaviour drifts. Data drifts faster. An agent that was accurate at launch and unmeasured since is an agent whose current accuracy nobody knows.

  6. 06

    Change control on the parts that actually change

    Model version, system prompt, tool definitions, data sources and retrieval configuration are all versioned, and a change to any of them is re-evaluated before it reaches a user. Failures of this kind are rarely breakages. Something was changed, by someone reasonable, without a regression suite to catch what the change cost.

  7. 07

    Disclosure by default

    Where a person is dealing with an agent, they are told. Where output is synthetic, it is marked. The Regulation already requires this of certain systems and of certain deployments; we treat it as the floor rather than the target.

Request a 30-minute diagnostic

Every engagement starts with the audit

How the work runs

Every automation engagement starts where every Novus Point engagement starts — with the audit. A 30-minute diagnostic at no charge, then a written position, then engagement if it is warranted. The method below picks up from that written position. Nothing is built before there is one.

  1. 01

    Scope against the written position

    We take the position document and turn it into a build brief: the process boundary, the systems touched, the data that enters and leaves, the decisions the agent may own, the oversight points, and the acceptance test. Regulatory classification is addressed here, not later — including the question of which role you occupy under the Regulation, because commissioning a bespoke system can place you differently from buying a product off a shelf, and the obligations attached to a provider are not the obligations attached to a deployer. This step also names what we are not building.

  2. 02

    Build

    Small, working increments against the acceptance test, with the governance surface built in the same pass as the function — logging, versioning, oversight hooks, the stop. Model access sits behind an interface from day one, so that no part of the system assumes a particular provider will still be there, or still behave the same way, in a year. Where the work needs legal, security or data expertise, the firm brings in advisers who work under our direction and on our responsibility.

  3. 03

    Evaluate

    The agent is measured against the fixed test set before anyone outside the project sees it, including the failure cases: ambiguous input, missing data, hostile prompting, and the situations where the correct behaviour is to refuse and escalate. Results are written down. If the agent does not clear the bar for a class of input, that class is routed to a person and stays routed to a person until it does.

  4. 04

    Deploy

    Into a defined scope, with the fallback in place and the people who will operate it briefed on what it does, what it cannot do, and how to stop it. Deployment is the point at which the documentation has to be complete rather than intended — the system description, the runbook, the log schema and retention decision, the oversight assignments.

  5. 05

    Operate and monitor

    Live sampling against the evaluation standard, drift and cost monitoring, a defined review cadence, and change control on every subsequent modification. Where a client wants the firm to run this, we run it. Where a client wants their own team to hold it, handover includes the runbook, the evaluation set and a period working alongside the people who will own it. Either is fine. Handing over a system with no means of measuring it is not.

Why “governed” is not optional

The Regulation does not care that it worked in the demo.

Regulation (EU) 2024/1689 has been in force since 1 August 2024, and it is already operative in the places that matter most to anyone deploying agents. It was amended in July 2026 by Regulation (EU) 2026/1744, the Digital Omnibus on AI, published in the Official Journal on 24 July 2026 and in force since 27 July 2026. Both instruments are cited below by article, because the amendments moved dates and, in one case, changed a duty.

  1. Article 2 — territorial scope

    01

    First question: does it reach you at all?

    This is an EU regulation, and a UK operation is not automatically inside it. Article 2 catches you in two situations that matter in practice: where you place an AI system on the Union market or put one into service there — as a provider — irrespective of where you are established; and where you are a provider or deployer established outside the Union but the output produced by the system is used in the Union. A London firm running an internal agent on UK-only data and UK-only users is, on the face of it, outside. A London firm whose assistant answers customers in Dublin, or whose agent produces output relied on by an EU subsidiary, is not. That determination is the single most valuable half-hour we spend with a new client, and it comes before any discussion of what to build. The rest of this section assumes the answer is yes.

  2. Applied 2 February 2025

    02

    The prohibitions are live.

    The Article 5 prohibitions have applied since 2 February 2025. Article 5 sets out eight prohibited practices, rising to ten on 2 December 2026, when the two added by Regulation (EU) 2026/1744 take effect — and clearing them is the first thing any build has to do.

  3. Applied 2 February 2025 · rewritten 27 July 2026

    03

    AI literacy is an obligation about your people — and its shape changed on 27 July 2026.

    Article 4 has applied since the same date, 2 February 2025, and it attaches to providers and deployers alike, reaching the staff and other persons dealing with the operation and use of the system on your behalf. Regulation (EU) 2026/1744 rewrote it with effect from 27 July 2026. The duty is now to take measures to support the development of AI literacy, and the amended text says expressly that it does not require providers or deployers to guarantee any specific level of AI literacy of any individual. That is an obligation of means rather than of result. It is lighter than it was. It is not satisfied by software, and it is still routinely missed.

  4. Enforceable 2 August 2025

    04

    Penalties are enforceable.

    Article 99 has been enforceable since 2 August 2025 — up to EUR 35m or 7% of global annual turnover for prohibited practices, and up to EUR 15m or 3% for most other breaches. The one carve-out is Article 101, the fines on providers of general-purpose AI models, which sits on its own timetable. Obligations for general-purpose AI models have applied since 2 August 2025.

  5. Applied 2 August 2026

    05

    Transparency is the obligation already in force.

    Article 50 has applied since 2 August 2026, and the Digital Omnibus did not defer it. It divides along the provider and deployer line, which is where most readers of this page misplace themselves:

    Article 50 — which side of the line you are on

    • As a provider Article 50(1): AI systems intended to interact directly with natural persons must be designed and developed so that those persons are informed they are dealing with an AI system, unless that is obvious to a reasonably well-informed person. Article 50(2): providers of systems generating synthetic audio, image, video or text — general-purpose systems included — must ensure the outputs are marked in a machine-readable format and detectable as artificially generated or manipulated. If you commission a bespoke assistant, you may well be the provider of it.
    • As a deployer Article 50(3): deployers of an emotion recognition system or a biometric categorisation system must inform the people exposed to it. Article 50(4): deployers must disclose deep fakes, and must disclose AI-generated or AI-manipulated text published to inform the public on matters of public interest, subject to narrow exceptions where a human holds editorial responsibility.

    There is one transitional. Generative systems already placed on the market before 2 August 2026 have until 2 December 2026 to bring their outputs into line with the machine-readable marking duty. Anything placed on the market on or after 2 August 2026 — which is to say anything we build for you — gets no such relief, and has to ship marked.

  6. From 2 December 2027 · Annex I from 2 August 2028

    06

    High-risk arrives on 2 December 2027, and that is not far enough off to build carelessly.

    Obligations for the high-risk systems listed in Annex III apply from 2 December 2027, and for high-risk AI embedded in the regulated products of Annex I from 2 August 2028, following the deferral in Regulation (EU) 2026/1744. If an agent you deploy lands in that category, Article 26 sets out what being the deployer costs: use in accordance with the instructions for use; human oversight assigned to natural persons with the necessary competence, training and authority; input data that is relevant and sufficiently representative where you control it; monitoring of the system in operation, with a duty to suspend use and inform the provider and the market surveillance authority where it presents a risk; logs kept for at least six months; where the deployer is an employer, workers’ representatives and affected workers informed before the system goes into service at the workplace; and, for the Annex III systems that make or assist decisions about people, those people told they are subject to it. That is not a document you produce in the final quarter. It is a description of how the system has to have been built.

    Read that list again and notice what it actually asks for: know what the system does, keep a record, put a competent human in the right place, tell people the truth, and be able to stop. None of it is exotic. It is close to a description of a well-engineered system. The organisations in difficulty are not the ones that found the Act demanding. They are the ones that built first and asked afterwards.

Our house position. We will not promise you compliance. The regulator decides that, not us. What we give you is an accurate picture and a defensible order of work — and, where we build, a system whose behaviour can actually be evidenced.

Request a 30-minute diagnostic

Four situations

Who this is for

01

The COO with a process that runs on people remembering things

The work is high volume, rule-heavy and interruption-driven. It is documented nowhere except in the heads of four people, two of whom are leaving. You do not need a strategy deck about AI. You need one process taken end to end, instrumented, and made to run the same way every time — and you need to be able to show your auditors how it decides.

02

The founder who bought a demo that died in production

It was extraordinary in the pitch. Six weeks in, it hallucinated in front of a customer, nobody could reconstruct what it had been asked, and it was quietly turned off. The system was never the problem in isolation. The absence of evaluation, logging and a defined escalation was. That is recoverable, and usually cheaper to fix than to start again — but it starts with an honest assessment of what is there.

03

The firm blocked by its own IT and security function

The business wants to adopt. Security will not sign. Both are behaving correctly: the objection is almost always about data flow, retention, access and the lack of an audit trail, and none of those are answered by enthusiasm. They are architecture questions with architecture answers, and they are far easier to answer at design stage than in a review board.

04

The board that has to sign something it does not fully understand

Directors are being asked to approve AI deployment while carrying personal exposure for governance they cannot inspect. What is needed is not reassurance. It is an inventory, a classification, a written position, and systems built so that the answers to the obvious questions already exist.

Five questions

Questions the diagnostic usually starts with

  1. 01

    Should we build or buy?

    Buy, where the process is generic and a mature product already fits it. Build, where the process is the thing that actually differentiates you, where the data cannot leave your control, or where every available product would require you to bend the business to the software. We will tell you to buy when buying is right, and we would rather say it in the diagnostic than in a proposal.

  2. 02

    Which models do you use?

    Whichever is right for the task, and the honest answer is that it will change. We do not have a house model and we take no position on which laboratory wins. What we do have is a rule: model access sits behind an interface, so the choice is a configuration decision rather than a structural commitment. Selection is recorded — which model, which version, why, and what it was measured against — so that the decision can be revisited on evidence rather than on preference. Where data residency or confidentiality is a constraint, we design so that the workload can run on infrastructure you control, and we say at scoping stage what that costs and what it rules out.

  3. 03

    What happens when the model provider changes the model?

    They will. Versions are deprecated, and behaviour shifts under a name that has not changed. This is a build risk, and it is managed like one. Versions are pinned where the provider allows pinning. Prompts, tool definitions and model versions live in version control. The evaluation set is the regression suite, so a provider-side change is caught by measurement rather than by a user. And because model access is abstracted, moving is an engineering task with a known cost, not a rebuild.

  4. 04

    Who owns the code?

    Ownership is settled in the engagement letter before work starts, and our default position is that you own it — source, documentation, evaluation set and runbook. Anything you would need in order to keep running or hand on the system is yours by default, and where a third-party component carries a licence obligation it is listed with that obligation, so nothing arrives with a condition you did not know you accepted.

  5. 05

    Our security team will not approve an external AI tool. Does that end the conversation?

    Usually it starts a better one. Security’s objection is rarely to AI as such; it is to unclear data flow, undefined retention, broad access and an absent audit trail. Those are the same things governance-by-design produces, which is why security review tends to go faster on a build designed this way than on a product bought this quarter. We would far rather have that conversation at design stage, with your security and data people in the room, than discover the objection after the money is spent.

Next step

Thirty minutes. Then you will know whether this is worth building.

What the diagnostic is

  • The diagnostic is a conversation, not a pitch. We ask what you run, who touches it, what you have already committed to, and where the process actually hurts. You get a straight answer on the call about whether an agent is the right instrument — and if it is not, we will say so.
  • If it is worth going further, you get a written position. What to build, in what order, what it will be measured against, and what it will cost to run once it exists. Engagement follows only if that document justifies it.

Request a 30-minute diagnostic

jakub@novus-point.co.uk — enquiries go to the firm’s principal and are answered directly.