In 2024 a Hong Kong finance worker was tricked into transferring USD $25 million after attending a video call with what he believed were senior colleagues — every participant was an AI-generated deepfake. By 2026 the technology to clone a voice from 30 seconds of audio is freely available, and NZ businesses are now reporting confirmed cases of AI voice cloning being used in attempted CEO fraud and authorised push payment scams.
The defining shift is this: every assumption your finance team makes about voice authentication — "I recognise her voice", "It sounded like him on a slightly bad line" — is no longer reliable. The only defence is process, and process only works if staff are trained to follow it under pressure.
How AI voice cloning attacks work in NZ
The pattern reported to NZ banks and CERT NZ is consistent:
- Reconnaissance — the attacker identifies the CEO or CFO of a target business through LinkedIn, the company website, or a corporate podcast appearance.
- Voice harvesting — 30 to 60 seconds of clean voice audio is collected from a podcast, AGM recording, conference talk, or even an outgoing voicemail message.
- Voice clone generation — a synthetic voice model is built using commercial AI tools, often in under 10 minutes.
- Pretext call — the attacker calls a finance staff member or the CEO's PA, often spoofing the executive's mobile number, with an urgent payment request: an acquisition that must close today, a regulator settlement, a supplier in crisis.
- Pressure and isolation — the script is designed to bypass normal verification: "I'm in a meeting", "Can't email — keep this off the system until it's done", "I'll explain when I'm back."
The audio quality is now indistinguishable from a real phone call for most listeners. Static, accent, hesitation, and emotion can all be modelled. Even the executive's catchphrases and verbal tics get picked up if there is enough source audio.
Why NZ executives are easier targets than US ones
Two structural factors make NZ businesses disproportionately vulnerable:
- Smaller leadership teams — in a 200-person NZ business, the CFO knows the AP clerk personally. That trust is the exploit.
- Public CEO profiles — NZ business culture rewards CEO visibility (industry conferences, media interviews, podcasts). That same visibility creates the voice corpus an attacker needs.
What staff training has to cover
Technical controls — caller ID, voiceprint authentication — are not reliable defences. Caller ID can be spoofed and voiceprint systems are being defeated by current-generation deepfakes. The only consistent defence is a callback policy with an authenticated channel that staff actually follow.
Effective training must cover:
- Mandatory callback to a known number — never the number that called in. The number must come from your authoritative directory, not from caller ID.
- Codeword / out-of-band verification — a pre-agreed phrase or shared secret used to verify identity on urgent payment requests. Rotated quarterly.
- Friction is the goal — staff must be empowered to delay payments and ask uncomfortable questions, even of the CEO. The CEO must publicly endorse this and be seen to follow the process themselves.
- Recognise the pressure pattern — urgency, secrecy, isolation, and authority are the four hallmarks. Any one of them should trigger callback.
This is the same control structure used to defend against business email compromise — the threat vector has moved from email to voice, but the human-process defence is the same.
The compliance picture
For organisations subject to APRA CPS 234 (relevant to NZ subsidiaries of Australian financial institutions) or NZ banking sector cyber requirements, fraud loss from authorised push payments now sits squarely inside operational risk and information security accountability. Boards expect documented training on synthetic media threats. Cyber insurers increasingly ask whether voice-based social engineering is covered in awareness programmes — a "no" can affect coverage and premiums.
For NZ government agencies, NZISM Control 3.2.18 requires awareness training that reflects current threats. AI-generated impersonation is the highest-velocity threat in 2026; static training programmes designed for 2022 do not satisfy the control's intent.
| Defence | Effective against AI voice cloning? |
|---|---|
| Caller ID verification | No — easily spoofed |
| Voice recognition by staff | No — synthetic voices are now indistinguishable |
| Voiceprint biometrics | Diminishing — being defeated by current-gen deepfakes |
| Mandatory callback to authoritative number | Yes |
| Out-of-band codeword verification | Yes |
| Trained staff who follow process under pressure | Yes |
The training component SecureAZ provides
SecureAZ awareness modules include synthetic media threat scenarios — AI voice clones, deepfake video calls, and the social engineering patterns that surround them. The simulations are NZ-contextualised: a CFO call from a "supplier in Auckland", an "urgent transfer to Sydney", a "regulator settlement" framed in NZ language and currency.
Combined with documented phishing training for employees and a callback policy your leadership team actually models, you reduce one of the highest-impact and lowest-volume threats in 2026 to a manageable risk.
Start a free SecureAZ trial — 45 days, NZ content, full compliance documentation.
External references: