Synthetic Australian Financial Planning Datasets — Complete Guide (2026)
Key takeaways:
- Synthetic only — no real client PII, not personal advice, for AI model training and research
- Five specialised types: client profiles, discovery conversations, risk scenarios, SOA-style content, and compliance judgements (s961B/s961G)
- Privacy-safe AU generation — s961B/RG175 grounded, no real PII
- Available as singles or all-inclusive bundle — pricing coming soon, enquire via contact
Synthetic data notice: This dataset is artificially generated for training and research. It does not contain real client information and is not personal financial advice. Use for model development only.
What Is Synthetic Financial Planning Data?
Synthetic financial planning data is artificially created to mirror real Australian advice scenarios without using real client records. Each record is a DatasetRecord with a payload (e.g. a client fact-find or SOA, ROA or CAR) and rich metadata (model, temp, s961B cues). It lets fintechs and AI labs train AU-specific models for discovery chat, risk profiling and SOA compliance without privacy risk or 12-month collection. For what good advice documents must contain, see our RG 175 guide and RG 90 example.
Unlike generic US datasets, AU synthetic must handle preservation age, concessional $32,500 and non-concessional $130,000 caps (from July 2026), HECS/HELP, Centrelink deeming (0.25% to $60,400 then 2.25%), Division 293/296, FHSS $15k, SMSF LRBA ban from 1 July 2024, and RG175 clear concise effective duties.
The Five Dataset Types
1. Client profiles — age, employment, income, super balance, SMSF flag, mortgage, family situation, goals, risk tolerance (conservative to high_growth), advice scope (full/scaled/limited). Used to simulate diverse AU clients.
2. Discovery conversations — 5-turn adviser–client dialogues with intents (fact_find, risk_profiling, compliance) and s961B safe harbour notes.
3. Risk scenarios — questionnaire score, capacity for loss, recommended portfolio (30/70 to 90/10) and conflicts.
4. SOA-style content — scope, strategy (salary sacrifice, debt recycling, TTR), fee disclosure ($3,300/$4,700), risk section and s961B/G compliance statement.
5. Compliance judgements — s961B/s961G/fee booleans, reasoning and label (compliant/non_compliant/borderline) for classifier training.
All five are linked by client_id → conversation→soa→compliance so a bundle of 25 = 125 lines that stay in sync.
How We Generate s961B-Compliant Data
We use an AU prompt library covering s961B, RG175, super, SMSF, risk, deeming and more, with privacy-safe synthetic generation. No real client data is used, and the pipeline is isolated so consumer trust is untouched.
Sample & Schema
Each line is JSONL — one record per line with payload and metadata. Truncated example:
{"record_type": "conversation", "payload": {"turns": [{"speaker": "adviser", "text": "Thanks for coming in..."}, {"speaker": "client", "text": "..."}]}}
Available as singles or all-inclusive bundle — enquire for sample.
Pricing — Coming Soon
Singles and bundle will be tiered by size — enquire now to join waitlist and lock early pricing. Stripe links to be published here.
Who Is It For?
WealthTech and RegTech teams building AU-specific compliance checks, advice chat, and risk tools. University researchers studying s961B and RG175, and enterprise AI labs needing privacy-safe AU training data without collecting real SOAs. If your model must understand Australian super, Centrelink and SMSF rules, synthetic gives you that grounding.
Banks and large financial planning firms with in-house models use the same datasets to fine-tune adviser copilots, paraplanner assistants and file-review QA, calibrate quality across advisers and licensees, train reviewers, and benchmark vendor tools — without exposing real client SOAs.
For In-House Model Teams at Banks and Licensees
For banks and AFSLs, real SOAs cannot be used for model training without creating privacy, confidentiality and accountability risk under APRA CPS 234 (information security), s912A general obligations (efficient, honest and fair services, compliance, risk management), and the Financial Accountability Regime (FAR) for accountable persons. Synthetic removes that blocker: 500–10k balanced compliant, borderline and non-compliant records, linked by client_id from fact-find through conversation and SOA to compliance judgement, ready for fine-tuning, regression testing and audit evidence.
Typical in-house uses: grounding internal copilots in s961B safe-harbour steps, RG175 clear-concise-effective duties and AFCA decision patterns; calibrating file reviewers across teams with shared edge cases (Division 296, deeming, LRBA ban, scaled vs full scope); onboarding paraplanners and advisers with realistic but fake files; and independently validating vendor AI before enterprise rollout. Available to any buyer — same singles and bundle, same synthetic-only, not-advice terms.
Why Synthetic Beats Real for AU Training
Real SOAs are protected by privacy and confidentiality obligations and cannot be shared; synthetic gives you 500–10k balanced compliant/non labels in days, not months, with no PII, no AFCA risk, and full control over edge cases (Division 296, deeming, LRBA ban). Your model learns AU law, not US 401k. You control the mix of scaled vs full advice, conservative vs high growth, and compliant vs non-compliant examples.
How to Use It
Load the JSONL into Python, Hugging Face datasets or your training pipeline — one line = one record. Use client profiles to simulate onboarding, conversations to fine-tune chat, and compliance labels to evaluate your checker. All records are synthetic, balanced and linked by client_id so you can trace a fact-find through to the SOA and the judgement.
How This Powers AdviserCheck
The same 6-layer compliance logic that checks your SOA trains these datasets — completeness, s961B, contradictions, s961G, language, fees. Synthetic lets us test at scale without touching real client data. Check your own SOA free while we build the data.
Related guides:
Frequently Asked Questions
Is this real client data?
No. 100% synthetic, artificially generated from AU prompt seeds. No real PII, not personal advice, for training/research only.
Can I buy one type or must I buy the bundle?
Both — 5 singles (e.g. just conversations) or all-inclusive bundle `rich_seed.jsonl` (5× size). Bundle is ~40% cheaper per line.
How many records do I need?
Starter 500 per file (500 rich) is enough to fine-tune, Pro 2k for evaluation, Enterprise 10k for production. We generate 25 per run, append to any size.
Will you add more variation?
Yes — generation uses varied AU scenarios and per-bundle personalisation so each record is unique.
Can banks use this for in-house models without using real SOAs?
Yes — designed for that. Fine-tune copilots, QA checkers and training tools on synthetic s961B/RG175 cases with no client PII, supporting CPS 234, s912A and FAR accountability expectations. Same datasets, open to any buyer.
Want sample or early pricing?
Enquire About DatasetsLast updated: 2026-09-14. Synthetic data — not personal advice. For model training and research.