10
Development teams should be staffed with qualified personnel with appropriate AI/ML
expertise and secure coding training. Clear roles and responsibilities should be defined among
data scientists, developers, validators, and business owners.
4.3 Data Controls.
The increasing importance of large, high-quality datasets for AI model training and operation
means that risk groups must become more involved in data controls and integrity than ever
before. A robust data strategy is a prerequisite for effective AI deployment. Firms must ensure
that data quality, lineage, and integrity are actively managed, as poor data can lead to flawed
model outputs and decisions. Data governance should be embedded into AI Systems from the
outset to support reliable performance and regulatory compliance.
• AI-Ready Data Characteristics
High-quality AI data needs to be complete (sufficient volume and coverage), accurate
(verified against trusted sources), timely (appropriate freshness for use case), consistent
(standardized formats and definitions), representative (capturing full range of market
conditions including tail events), and traceable (clear lineage from source to
consumption).
• Data Governance Framework
Data quality management should include automated quality checks, anomaly detection,
and data profiling. Data lineage tracking should be in place and provide full
documentation of data flows from source through transformation to model input. Role-
based access controls should align with data sensitivity. Privacy protections are essential
for confidential market positions and proprietary information, and therefore considered
during data control processes. Vendor data management due diligence should occur for
SLAs and contingency plans should be created for third-party providers.
4.4 Validation
As with any model, AI Tools should undergo a rigorous validation process prior to moving to
production. Validation should be proportional to risk tier and use case. For low-risk tasks (e.g.,
email drafting), simple test cases may suffice. For high-risk tasks (e.g., automated transactions
and confirmations), rigorous validation is needed. Validation should include stress testing,
qualitative assessments, data time horizon sensitivity, and independent challenge.
Investigation of edge cases, especially performance in tail-risk scenarios, is another key
validation point for market-centered tools.
AI specific validation can include methods like train/test splits, cross validation, bootstrapping,
LLM fact-checking, and statistical testing (e.g., sMAPE, MAE, and MSE). Continuous learning
models should be subject to ongoing monitoring with alerts for threshold breaches (e.g. position
limits).
Validators should apply effective challenge principles and document their findings to ensure
model integrity and accountability. Validators should note sensitivities and limitations. They
should also develop a safety checklist for deployment, recommended initial limits, and clear
guidance for manual checks or guardrails for eventual usage.
Development teams should be staffed with qualified personnel with appropriate AI/ML
expertise and secure coding training. Clear roles and responsibilities should be defined among
data scientists, developers, validators, and business owners.
4.3 Data Controls.
The increasing importance of large, high-quality datasets for AI model training and operation
means that risk groups must become more involved in data controls and integrity than ever
before. A robust data strategy is a prerequisite for effective AI deployment. Firms must ensure
that data quality, lineage, and integrity are actively managed, as poor data can lead to flawed
model outputs and decisions. Data governance should be embedded into AI Systems from the
outset to support reliable performance and regulatory compliance.
• AI-Ready Data Characteristics
High-quality AI data needs to be complete (sufficient volume and coverage), accurate
(verified against trusted sources), timely (appropriate freshness for use case), consistent
(standardized formats and definitions), representative (capturing full range of market
conditions including tail events), and traceable (clear lineage from source to
consumption).
• Data Governance Framework
Data quality management should include automated quality checks, anomaly detection,
and data profiling. Data lineage tracking should be in place and provide full
documentation of data flows from source through transformation to model input. Role-
based access controls should align with data sensitivity. Privacy protections are essential
for confidential market positions and proprietary information, and therefore considered
during data control processes. Vendor data management due diligence should occur for
SLAs and contingency plans should be created for third-party providers.
4.4 Validation
As with any model, AI Tools should undergo a rigorous validation process prior to moving to
production. Validation should be proportional to risk tier and use case. For low-risk tasks (e.g.,
email drafting), simple test cases may suffice. For high-risk tasks (e.g., automated transactions
and confirmations), rigorous validation is needed. Validation should include stress testing,
qualitative assessments, data time horizon sensitivity, and independent challenge.
Investigation of edge cases, especially performance in tail-risk scenarios, is another key
validation point for market-centered tools.
AI specific validation can include methods like train/test splits, cross validation, bootstrapping,
LLM fact-checking, and statistical testing (e.g., sMAPE, MAE, and MSE). Continuous learning
models should be subject to ongoing monitoring with alerts for threshold breaches (e.g. position
limits).
Validators should apply effective challenge principles and document their findings to ensure
model integrity and accountability. Validators should note sensitivities and limitations. They
should also develop a safety checklist for deployment, recommended initial limits, and clear
guidance for manual checks or guardrails for eventual usage.

















