NIST's AI Metrology Center is a set of resources within the AI Resource Center (AIRC) that supports the broad adoption of trustworthy AI. The Center integrates metrics, methodologies and tools with the NIST AI RMF trustworthy characteristics and lifecycle stages, enabling organizations to identify and use the most appropriate methodologies for their AI use cases. Organizations are encouraged to select measurement approaches that can strengthen their own testing, evaluation, validation, and verification (TEVV) and governance processes, facilitating more reliable deployment and iteration for AI systems. Inclusion of resources in the Center does not constitute NIST endorsement, validation, or determination of suitability of any listed methodology, metric, or tool.

Community submissions

How submissions work

Community submissions of AI metrics and measurement methodologies are open. Proposals for inclusion in the Center are accepted and reviewed in the open on GitHub: a submission is a pull request that adds one YAML file describing a single metric or measurement method, review happens publicly on that pull request, and an accepted submission is considered for publication here.

The usnistgov/ai-metrology-submissions repository holds the submission guide, the authoritative format specification, and worked example submissions kept open so you can read a real review end to end.

Submit your metric for consideration

This is a new process and we expect to refine it as we go. Do not submit proprietary or confidential information — submission files and all review discussion are publicly visible.

AI Metrics and Measurement Methods

  • Metric
    Name
    Adversariality 
    Applied Definition
    Percentage of randomly sampled viewpoints classified as the adversarial class.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    1ref
    • Anish Athalye, Logan Engstrom, Andrew Ilyas, and Kevin Kwok. Synthesizing Robust Adversarial Examples. In Proceedings of the 35th International Conference on Machine Learning, pages 284–293. PMLR, July 2018. ISSN: 2640-3498. Link
    Tools
    2tools
  • Measurement Method
    Name
    Agent / Tool Abuse Testing 
    Applied Definition
    Testing whether a system misuses connected tools or external actions through, e.g., unsafe tool selection, excessive agency, unauthorized action attempts, or harmful task execution.
    AI RMF Characteristic(s)
    Valid & ReliableSafeSecure & Resilient
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Red Teaming Evaluation Method
    References
    4refs
    • Bryan, Pete, Giorgio Severi, Joris de Gruyter, Daniel Jones, Blake Bullwinkel, Amanda Minnich, Shiven Chawla et al. "Taxonomy of failure mode[s] in agentic AI systems." Microsoft AI Red Team (2025)
    • Shi, Jiawen, Zenghui Yuan, Guiyao Tie, Pan Zhou, Neil Zhenqiang Gong, and Lichao Sun. "Prompt injection attack to tool selection in LLM agents." arXiv preprint arXiv:2504.19793 (2025)
    • OWASP LLM05:2025 Improper Output Handling, LLM06:2025 Excessive Agency, LLM10:2025 Unbounded Consumption
    • MITRE ATLAS AI Agent Tool Invocation, Exfiltration via AI Agent Tool Invocation, AI Agent Tool Poisoning, Publish Poisoned AI Agent Tool, Cost Harvesting.
    Tools
    9tools
  • Measurement Method
    Name
    ARC AGI 
    Applied Definition
    A benchmark series with focus on reasoning; offers public tasks and private hold-out data for official evaluations.
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Chollet, Francois, Mike Knoop, Gregory Kamradt, Bryan Landers, and Henry Pinkard. "ARC-AGI-2: A new challenge for frontier AI reasoning systems." arXiv preprint arXiv:2505.11831 (2025).
  • Metric
    Name
    Attack Success Rate (AI Privacy) 
    Applied Definition
    The fraction of test instances for which an adversary correctly associates a subject's identity or sensitive attributes from an AI model or dataset.
    AI RMF Characteristic(s)
    Privacy-Enhanced
    AI Lifecycle Stage(s)
    BuildUseDeployOperate & Monitor
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    2refs
    • Ashwin Machanavajjhala, Daniel Kifer, Johannes Gehrke, and Muthuramakrishnan Venkitasubramaniam. L-diversity: Privacy beyond k-anonymity. ACM Transactions on Knowledge Discovery from Data, 1(1):3–es, March 2007. ISSN 1556-4681. doi:10.1145/1217299.1217302. Link
    • Runhua Xu, Nathalie Baracaldo, and James Joshi. Privacy-Preserving Machine Learning: Methods, Challenges and Directions. Technical Report arXiv:2108.04417, arXiv, September 2021. arXiv:2108.04417 [cs] type: article. Link
  • Metric
    Name
    Attack Success Rate (AI Security) 
    Applied Definition
    Percentage of generated adversarial inputs that are misclassified.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    2refs
    • Pin-Yu Chen, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh. ZOO: Zeroth Order Optimization Based Black-box Attacks to Deep Neural Networks without Training Substitute Models. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security, AISec ’17, pages 15–26, New York, NY, USA, November 2017. Association for Computing Machinery. ISBN 978-1-4503-5202-4. doi: 10.1145/3128572.3140448. Link
    • Noam Yefet, Uri Alon, and Eran Yahav. Adversarial examples for models of code. Proceedings of the ACM on Programming Languages, 4(OOPSLA):162:1–162:30, November 2020. doi: 10.1145/3428230. . Link
  • Measurement Method
    Name
    Availability Attacks 
    Applied Definition
    Testing whether a system can be degraded, delayed, or access to it can be denied through, e.g., sponge examples attacks, repeated token attacks, or input overflows.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Red Teaming Evaluation Method
    References
    4refs
    • Shumailov, Ilia, Yiren Zhao, Daniel Bates, Nicolas Papernot, Robert Mullins, and Ross Anderson. "Sponge examples: Energy-latency attacks on neural networks." In 2021 IEEE European symposium on security and privacy (EuroS&P), pp. 212-231. IEEE, 2021
    • Yona, Itay, Ilia Shumailov, Jamie Hayes, and Yossi Gandelsman. "Interpreting the Repeated Token Phenomenon in Large Language Models." In Forty-second International Conference on Machine Learning
    • OWASP LLM10:2025 Unbounded Consumption
    • MITRE ATLAS Denial of AI Service.
    Tools
    8tools
  • Metric
    Name
    Average Input Change Amount 
    Applied Definition
    Mean value of input modification across an adversarial set.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    1ref
    • Jiliang Zhang and Chen Li. Adversarial Examples: Opportunities and Challenges. IEEE Transactions on Neural Networks and Learning Systems, 31(7):2578–2593, July 2020. ISSN 2162-2388. doi: 10.1109/TNNLS.2019.2933524. Conference Name: IEEE Transactions on Neural Networks and Learning Systems. Link
  • Metric
    Name
    Average Sensitivity 
    Applied Definition
    Mean change in explanation distance within a local neighborhood.
    AI RMF Characteristic(s)
    Valid & ReliableExplainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeployOperate & Monitor
    Primary TEVV Application
    Interpretability & Explanation Evaluation
    References
    1ref
    • Umang Bhatt, José M. F. Moura, and Adrian Weller. Evaluating and Aggregating Feature-based Model Explanations. volume 3, pages 3016–3022, July 2020. doi: 10.24963/ijcai.2020/417. ISSN: 1045-0823. Link
  • Measurement Method
    Name
    BELEBELE 
    Applied Definition
    Multilingual reading-comprehension benchmark designed to assess language coverage. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableFairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Bandarkar, Lucas, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. "The BELEBELE benchmark: a parallel reading comprehension dataset in 122 language variants." In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 749-775. 2024.
  • Measurement Method
    Name
    BFW 
    Applied Definition
    Facial recognition benchmark with labeled faces in realistic conditions and demographic labels. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableFairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Facial Recognition Benchmark Evaluation
    References
    1ref
    • Robinson, Joseph P., Gennady Livitz, Yann Henon, Can Qin, Yun Fu, and Samson Timoner. "Face recognition: too bias, or not too bias?." In Proceedings of the ieee/cvf conference on computer vision and pattern recognition workshops, 2020.
  • Measurement Method
    Name
    C-Eval 
    Applied Definition
    Chinese-language evaluation suite. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableFairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Huang, Yuzhen, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu et al. "C-EVAL: a multi-level multi-discipline Chinese evaluation suite for foundation models." In Proceedings of the 37th International Conference on Neural Information Processing Systems, pp. 62991-63010. 2023.
    Tools
    1tool
  • Metric
    Name
    Change in Performance Accuracy 
    Applied Definition
    Delta in system accuracy when moving from clean to adversarial inputs.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeployOperate & Monitor
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    1ref
    • Battista Biggio, Igino Corona, DavideMaiorca, Blaine Nelson, Nedim Srndic, Pavel Laskov, Giorgio Giacinto, and Fabio Roli. Evasion attacks against machine learning at test time. In Joint European conference on machine learning and knowledge discovery in databases, pages 387–402. Springer, 2013.
  • Measurement Method
    Name
    Change in Performance Accuracy of Human Output 
    Applied Definition
    Delta in human task accuracy with vs. without system explanations. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Explainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Human-Centered Evaluation
    References
    13refs
    • Andrew Anderson, Jonathan Dodge, Amrita Sadarangani, Zoe Juozapaitis, Evan Newman, Jed Irvine, Souti Chattopadhyay, Matthew Olson, Alan Fern, and Margaret Burnett. Mental Models of Mere Mortals with Explanations of Reinforcement Learning. ACM Transactions on Interactive Intelligent Systems, 10(2):15:1– 15:37, May 2020. ISSN 2160-6455. doi: 10.1145/3366485. Link
    • Been Kim, Cynthia Rudin, and Julie A Shah. The Bayesian Case Model: A Generative Approach for Case-Based Reasoning and Prototype Classification. In Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, and K. Q. Weinberger, editors, Advances in Neural Information Processing Systems 27, pages 1952–1960. Curran Associates, Inc., 2014. Link
    • Isaac Lage, Emily Chen, Jeffrey He,Menaka Narayanan, Been Kim, Sam Gershman, and Finale Doshi-Velez. An Evaluation of the Human-Interpretability of Explanation. arXiv:1902.00006 [cs, stat], August 2019. arXiv: 1902.00006 Link
    • Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Samuel J. Gershman, and Finale Doshi-Velez. Human Evaluation of Models Built for Interpretability. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, 7(1):59–67, October 2019. Link
    • Vivian Lai and Chenhao Tan. On Human Predictions with Explanations and Predictions of Machine Learning Models: A Case Study on Deception Detection. arXiv, November 2018. doi: 10.1145/3287560.3287590. Link
    • Himabindu Lakkaraju, Ece Kamar, Rich Caruana, and Jure Leskovec. Faithful and Customizable Explanations of Black Box Models. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, AIES ’19, pages 131–138, New York, NY, USA, 2019. Association for Computing Machinery. ISBN 978-1-4503-6324-2. doi: 10.1145/3306618.3314229. . event-place: Honolulu, HI, USA Link
    • Oisin Mac Aodha, Shihan Su, Yuxin Chen, Pietro Perona, and Yisong Yue. Teaching Categories to Human Learners with Visual Explanations. arXiv, February 2018. Link
    • Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer Wortman Vaughan, and Hanna Wallach. Manipulating and Measuring Model Interpretability. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, CHI ’21, pages 1–52, New York, NY, USA, May 2021. Association for Computing Machinery. ISBN 978-1-4503-8096-6. doi: 10.1145/3411764.3445315. Link
    • Forough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan, and Hanna Wallach. Manipulating and Measuring Model Interpretability. arXiv:1802.07810 [cs], November 2019. arXiv: 1802.07810 Link
    • Philipp Schmidt and Felix Biessmann. Quantifying Interpretability and Trust in Machine Learning Systems. arXiv:1901.08558 [cs, stat], January 2019. arXiv: 1901.08558 Link
    • Peter Hase and Mohit Bansal. Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5540–5552, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.491. Link
    • Dylan Slack, Sorelle A. Friedler, Carlos Scheidegger, and Chitradeep Dutta Roy. Assessing the Local Interpretability of Machine Learning Models. arXiv:1902.03501 [cs, stat], August 2019. arXiv: 1902.03501 Link
    • Yaniv Yacoby, Ben Green, Christopher L. Griffin, and Finale Doshi Velez. ”If it didn’t happen, why would I change my decision?”: How Judges Respond to Counterfactual Explanations for the Public Safety Assessment. Technical Report arXiv:2205.05424, arXiv, May 2022. arXiv:2205.05424 [cs] type: article Link
  • Metric
    Name
    Cohen's d 
    Applied Definition
    A standardized measure of effect size mapping the difference between two group means: d = (x1_bar - x2_bar) / s_pooled. Quantifies the magnitude of disparate impact or feature distribution shifts between groups. Associated thresholds are +/- 0.2, +/-0.5, +/-0.8 for small, medium, and large differences, respectively.
    AI RMF Characteristic(s)
    Valid & ReliableFair
    AI Lifecycle Stage(s)
    Collect & Process DataBuildUseOperate & Monitor
    Primary TEVV Application
    Fairness Evaluation
    References
    1ref
    • Cohen, Jacob. 2013. Statistical Power Analysis for the Behavioral Sciences. Routledge.
  • Metric
    Name
    Cohen's Kappa 
    Applied Definition
    A metric evaluating inter-annotator agreement for categorical items between exactly two annotators, correcting for chance: kappa = (p_o - p_e) / (1 - p_e).
    AI RMF Characteristic(s)
    Valid & ReliableAccountable & Transparent
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Fairness Evaluation
    References
    1ref
    • McHugh, Mary L. 2012. “Interrater Reliability: The Kappa Statistic.” Biochemia Medica 22 (3): 276–82. Link
  • Metric
    Name
    Compliance Rate 
    Applied Definition
    Percentage of input prompts that a model complies with.
    AI RMF Characteristic(s)
    Valid & ReliableSafe
    AI Lifecycle Stage(s)
    Collect & Process DataBuildUseOperate & Monitor
    Primary TEVV Application
    Safety Evaluation
    References
    1ref
    • Brahman, Faeze, Sachin Kumar, Vidhisha Balachandran, et al. 2024. “The Art of Saying No: Contextual Noncompliance in Language Models.” In Advances in Neural Information Processing Systems, edited by A. Globerson, L. Mackey, D. Belgrave, et al., vol. 37. Curran Associates, Inc. Link
  • Metric
    Name
    Conditional Statistical Parity 
    Applied Definition
    Predictions Y^, a set of features X, and group A should have equal probability (P) of a positive outcome: P(Y^|X=1,A=0)=P(Y^|X=1,A=1).
    AI RMF Characteristic(s)
    Valid & ReliableFair
    AI Lifecycle Stage(s)
    Collect & Process DataBuildUseOperate & Monitor
    Primary TEVV Application
    Fairness Evaluation
    References
    1ref
    • Mehrabi, Ninareh, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. “A Survey on Bias and Fairness in Machine Learning.” ACM Computing Surveys (CSUR) 54 (6): 1–35. Link
  • Measurement Method
    Name
    Confidentiality Attacks 
    Applied Definition
    Testing whether a system exposes confidential information or internal functionality through, e.g., autocompletion, eliciting or inferring user information, model extraction, model inversion, membership inference, or system prompt leakage/instruction disclosure.
    AI RMF Characteristic(s)
    Secure & ResilientPrivacy-Enhanced
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Red Teaming Evaluation Method
    References
    6refs
    • Barreno, Marco, Blaine Nelson, Anthony D. Joseph, and J. Doug Tygar. "The security of machine learning." Machine learning 81, no. 2 (2010): 121-148
    • Tramèr, Florian, Fan Zhang, Ari Juels, Michael K. Reiter, and Thomas Ristenpart. "Stealing machine learning models via prediction APIs." In 25th USENIX security symposium (USENIX Security 16), pp. 601-618. 2016
    • Duan, Michael, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettlemoyer, Yulia Tsvetkov, Yejin Choi, David Evans, and Hannaneh Hajishirzi. "Do Membership Inference Attacks Work on Large Language Models?." In First Conference on Language Modeling
    • OWASP LLM02:2025 Sensitive Information Disclosure
    • LLM07:2025 System Prompt Leakage
    • MITRE ATLAS LLM Data Leakage, Exfiltration.
    Tools
    8tools
  • Metric
    Name
    Counterfactual Fairness 
    Applied Definition
    For probabilities P, predictions Y^, a set of observed features X, a set of unobserved features U, and group A, a decision is fair if it remains the same in the actual world as in a counterfactual world where the individual belonged to a different demographic group: (P(Y^_A←a​(U)=y|X=x,A=a)=P(Y^_A←a′​(U)=y|X=x,A=a).
    AI RMF Characteristic(s)
    Valid & ReliableFair
    AI Lifecycle Stage(s)
    Collect & Process DataBuildUseOperate & Monitor
    Primary TEVV Application
    Fairness Evaluation
    References
    1ref
    • Mehrabi, Ninareh, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. “A Survey on Bias and Fairness in Machine Learning.” ACM Computing Surveys (CSUR) 54 (6): 1–35. Link
    Tools
    1tool
  • Measurement Method
    Name
    Counterfactual Fairness Prompting 
    Applied Definition
    Testing for disparate system outputs or outcomes by holding the task constant while varying demographic groups, languages, or dialects in counterfactual example prompts.
    AI RMF Characteristic(s)
    FairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Red Teaming Evaluation Method
    References
    2refs
    • NIST. "Artificial intelligence risk management framework: Generative artificial intelligence profile." NIST Trustworthy and Responsible AI Gaithersburg, MD, USA (2024)
    • Sturman, Olivia, Aparna R. Joshi, Bhaktipriya Radharapu, Piyush Kumar, and Renee Shelby. "Debiasing text safety classifiers through a fairness-aware ensemble." In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, pp. 199-214. 2024.
    Tools
    4tools
    • note-taking apps
  • Measurement Method
    Name
    CrowS-Pairs 
    Applied Definition
    Benchmark that measures stereotypical preferences across sentence pairs. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    FairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Nangia, Nikita, Clara Vania, Rasika Bhalerao, and Samuel Bowman. "CrowS-pairs: A challenge dataset for measuring social biases in masked language models." In Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), pp. 1953-1967. 2020.
  • Metric
    Name
    Decision Set Cardinality 
    Applied Definition
    Total number of rules in a decision set.
    AI RMF Characteristic(s)
    Explainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Interpretability & Explanation Evaluation
    References
    1ref
    • Himabindu Lakkaraju, Stephen H. Bach, and Jure Leskovec. Interpretable Decision Sets: A Joint Framework for Description and Prediction. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’16, pages 1675–1684, New York, NY, USA, 2016. Association for Computing Machinery. ISBN 978-1-4503-4232-2. doi: 10.1145/2939672.2939874. event-place: San Francisco, California, USA. Link
    Tools
    2tools
  • Metric
    Name
    Decision Set Coverage 
    Applied Definition
    Cardinality of the smallest rule set assigning a decision to every class.
    AI RMF Characteristic(s)
    Explainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Interpretability & Explanation Evaluation
    References
    1ref
    • Himabindu Lakkaraju, Stephen H. Bach, and Jure Leskovec. Interpretable Decision Sets: A Joint Framework for Description and Prediction. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’16, pages 1675–1684, New York, NY, USA, 2016. Association for Computing Machinery. ISBN 978-1-4503-4232-2. doi: 10.1145/2939672.2939874. event-place: San Francisco, California, USA. Link
    Tools
    2tools
  • Metric
    Name
    Decision Set Overlap 
    Applied Definition
    Count of data instances classified by more than one rule.
    AI RMF Characteristic(s)
    Explainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Interpretability & Explanation Evaluation
    References
    1ref
    • Himabindu Lakkaraju, Stephen H. Bach, and Jure Leskovec. Interpretable Decision Sets: A Joint Framework for Description and Prediction. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’16, pages 1675–1684, New York, NY, USA, 2016. Association for Computing Machinery. ISBN 978-1-4503-4232-2. doi: 10.1145/2939672.2939874. event-place: San Francisco, California, USA. Link
    Tools
    2tools
  • Measurement Method
    Name
    DecodingTrust: Fairness / Stereotype Bias / Toxicity 
    Applied Definition
    DecodingTrust components that assess fairness, stereotype bias, and toxicity. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableFairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Wang, Boxin, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al. "DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models." In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS '23), Article No. 1361, 31232–31339. Published May 30, 2024. Link
  • Metric
    Name
    Defense Efficacy 
    Applied Definition
    Fraction of adversarial inputs that are correctly classified by a defense.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    1ref
    • Yannik Potdevin, Dirk Nowotka, and Vijay Ganesh. An Empirical Investigation of Randomized Defenses against Adversarial Attacks, September 2019. . arXiv:1909.05580 [cs, stat]. Link
  • Metric
    Name
    Defense Quality 
    Applied Definition
    Accuracy on benign (non-adversarial) inputs after applying a defense.
    AI RMF Characteristic(s)
    Valid & ReliableSecure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeployOperate & Monitor
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    1ref
    • Yannik Potdevin, Dirk Nowotka, and Vijay Ganesh. An Empirical Investigation of Randomized Defenses against Adversarial Attacks, September 2019. arXiv:1909.05580 [cs, stat]. Link
  • Metric
    Name
    Defense Robustness 
    Applied Definition
    Fraction of random instantiations of a defense that remain effective.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    1ref
    • Yannik Potdevin, Dirk Nowotka, and Vijay Ganesh. An Empirical Investigation of Randomized Defenses against Adversarial Attacks, September 2019. arXiv:1909.05580 [cs, stat]. Link
  • Metric
    Name
    Demographic Parity 
    Applied Definition
    A metric testing that, for predictions Y^, a set of features X, and group A, the probability (P) of a positive outcome is the same for both across groups: (P(Y^=1|A=0)=P(Y^=1|A=1); common threshold of 0.8, or the "Four Fifth's Rule."
    AI RMF Characteristic(s)
    Valid & ReliableFair
    AI Lifecycle Stage(s)
    Collect & Process DataBuildUseOperate & Monitor
    Primary TEVV Application
    Fairness Evaluation
    References
    1ref
    • Mehrabi, Ninareh, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. “A Survey on Bias and Fairness in Machine Learning.” ACM Computing Surveys (CSUR) 54 (6): 1–35. Link
  • Metric
    Name
    Differential Privacy δ (delta) 
    Applied Definition
    Parameter that indicates the probability that the differential privacy guarantee fails to hold.
    AI RMF Characteristic(s)
    Privacy-Enhanced
    AI Lifecycle Stage(s)
    Collect & Process DataBuildUse
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    2refs
    • Cynthia Dwork. Differential Privacy: A Survey of Results. In Manindra Agrawal, Dingzhu Du, Zhenhua Duan, and Angsheng Li, editors, Theory and Applications of Models of Computation, number 4978 in Lecture Notes in Computer Science, pages 1–19. Springer Berlin Heidelberg, 2008. ISBN 978-3-540-79227-7 978-3- 540-79228-4. Link
    • Maoguo Gong, Yu Xie, Ke Pan, Kaiyuan Feng, and A.K. Qin. A Survey on Differentially Private Machine Learning [Review Article]. IEEE Computational Intelligence Magazine, 15(2):49–64, May 2020. ISSN 1556-6048. doi: 10.1109/MCI.2020.2976185. Conference Name: IEEE Computational Intelligence Magazine.
  • Metric
    Name
    Differential Privacy ε (epsilon) 
    Applied Definition
    Parameter for the privacy budget in differential privacy; lower is more private.
    AI RMF Characteristic(s)
    Privacy-Enhanced
    AI Lifecycle Stage(s)
    Collect & Process DataBuildUse
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    3refs
    • Cynthia Dwork. Differential Privacy: A Survey of Results. In Manindra Agrawal, Dingzhu Du, Zhenhua Duan, and Angsheng Li, editors, Theory and Applications of Models of Computation, number 4978 in Lecture Notes in Computer Science, pages 1–19. Springer Berlin Heidelberg, 2008. ISBN 978-3-540-79227-7 978-3- 540-79228-4. Link
    • Maoguo Gong, Yu Xie, Ke Pan, Kaiyuan Feng, and A.K. Qin. A Survey on Differentially Private Machine Learning [Review Article]. IEEE Computational Intelligence Magazine, 15(2):49–64, May 2020. ISSN 1556-6048. doi: 10.1109/MCI.2020.2976185. Conference Name: IEEE Computational Intelligence Magazine
    • Harsha Nori, Rich Caruana, Zhiqi Bu, Judy Hanwen Shen, and Janardhan Kulkarni. Accuracy, Interpretability, and Differential Privacy via Explainable Boosting. arXiv:2106.09680 [cs], June 2021. arXiv: 2106.09680. Link
  • Metric
    Name
    Discernibility Metric 
    Applied Definition
    Penalty score based on the number of indistinguishable tuples in a transformed dataset.
    AI RMF Characteristic(s)
    Privacy-Enhanced
    AI Lifecycle Stage(s)
    Collect & Process Data
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    2refs
    • R.J. Bayardo and Rakesh Agrawal. Data privacy through optimal k-anonymization. In 21st International Conference on Data Engineering (ICDE’05), pages 217–228, April 2005. doi: 10.1109/ICDE.2005.42. ISSN: 2375-026X
    • Ninghui Li, Tiancheng Li, and Suresh Venkatasubramanian. t-Closeness: Privacy Beyond k-Anonymity and l-Diversity. In 2007 IEEE 23rd International Conference on Data Engineering, pages 106–115, April 2007. doi: 10.1109/ICDE.2007.367856. ISSN: 2375-026X.
  • Metric
    Name
    Earth Mover's Distance (EMD) 
    Applied Definition
    Distance metric between probability distributions used for t-closeness.
    AI RMF Characteristic(s)
    Valid & ReliablePrivacy-Enhanced
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    1ref
    • Ninghui Li, Tiancheng Li, and Suresh Venkatasubramanian. t-Closeness: Privacy Beyond k-Anonymity and l-Diversity. In 2007 IEEE 23rd International Conference on Data Engineering, pages 106–115, April 2007. doi: 10.1109/ICDE.2007.367856. ISSN: 2375-026X.
  • Metric
    Name
    Encryption Overhead Ratio 
    Applied Definition
    Ratio of increase in the size of the ciphertext after encryption.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    Collect & Process Data
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    1ref
    • Abbas Acar, Hidayet Aksu, A. Selcuk Uluagac, and Mauro Conti. A Survey on Homomorphic Encryption Schemes: Theory and Implementation. ACM Computing Surveys, 51(4):79:1–79:35, July 2018. ISSN 0360-0300. doi: 10.1145/3214303. Link
    Tools
    3tools
  • Metric
    Name
    Equalized Odds 
    Applied Definition
    A metric that, for probabilities P, predictions Y^, and group A, tests whether protected and unprotected groups have equal rates for true positives and false positives: (P(Y^=1|A=0,Y=y)=P(Y^=1|A=1,Y=y) for y in {0, 1}.
    AI RMF Characteristic(s)
    Valid & ReliableFair
    AI Lifecycle Stage(s)
    Collect & Process DataBuildUseOperate & Monitor
    Primary TEVV Application
    Fairness Evaluation
    References
    1ref
    • Mehrabi, Ninareh, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. “A Survey on Bias and Fairness in Machine Learning.” ACM Computing Surveys (CSUR) 54 (6): 1–35. Link
  • Metric
    Name
    Equal Opportunity 
    Applied Definition
    A metric that, for predictions Y^, tests whether the probability of a person in the positive class (labels Y=1) being correctly assigned a positive outcome is the same for all groups A: P(Y^=1|A=0,Y=1)=P(Y^=1|A=1,Y=1).
    AI RMF Characteristic(s)
    Valid & ReliableFair
    AI Lifecycle Stage(s)
    Collect & Process DataBuildUseOperate & Monitor
    Primary TEVV Application
    Fairness Evaluation
    References
    1ref
    • Mehrabi, Ninareh, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. “A Survey on Bias and Fairness in Machine Learning.” ACM Computing Surveys (CSUR) 54 (6): 1–35. Link
  • Measurement Method
    Name
    ETHICS 
    Applied Definition
    Benchmark suite for evaluating ethics-related model outputs. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Accountable & Transparent
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Hendrycks, Dan, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. "Aligning AI With Shared Human Values." In International Conference on Learning Representations. 2021.
  • Metric
    Name
    Expectation Over Transformation (EOT) Distance 
    Applied Definition
    Expected distance of a transformed input from the original.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    1ref
    • Anish Athalye, Logan Engstrom, Andrew Ilyas, and Kevin Kwok. Synthesizing Robust Adversarial Examples. In Proceedings of the 35th International Conference on Machine Learning, pages 284–293. PMLR, July 2018. ISSN: 2640-3498. Link
  • Metric
    Name
    Explanation Correlation 
    Applied Definition
    A quantitative metric that measures the statistical relationship between two sets of explanations to determine if they change appropriately in response to altered inputs or model parameters.
    AI RMF Characteristic(s)
    Valid & ReliableExplainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Interpretability & Explanation Evaluation
    References
    2refs
    • Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. Sanity Checks for Saliency Maps. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31, pages 9505–9515. Curran Associates, Inc., 2018. Link
    • Leon Sixt, Maximilian Granz, and Tim Landgraf. When Explanations Lie: Why Many Modified BP Attributions Fail. arXiv, December 2019. Link
  • Measurement Method
    Name
    Explanation Satisfaction Scale 
    Applied Definition
    Multi-item scale measuring user satisfaction with explanations. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Explainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Human-Centered Evaluation
    References
    1ref
    • Robert R. Hoffman, Shane T. Mueller, Gary Klein, and Jordan Litman. Metrics for Explainable AI: Challenges and Prospects. arXiv:1812.04608 [cs], February 2019. arXiv: 1812.04608. Link
  • Metric
    Name
    Extraction Accuracy 
    Applied Definition
    Average accuracy (of the approximate model relative to the model) both on the test set and for a set of vectors uniformly chosen in the input space.
    AI RMF Characteristic(s)
    Privacy-Enhanced
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    2refs
    • Florian Tramer, Fan Zhang, Ari Juels, Michael K. Reiter, and Thomas Ristenpart. Stealing machine learning models via prediction APIs. In Proceedings of the 25th USENIX Conference on Security Symposium, SEC’16, pages 601–618, USA, August 2016. USENIX Association. ISBN 978-1-931971-32-4
    • Binghui Wang and Neil Zhenqiang Gong. Stealing Hyperparameters in Machine Learning. In 2018 IEEE Symposium on Security and Privacy (SP), pages 36–52, May 2018. doi: 10.1109/SP.2018.00038. ISSN: 2375-1207.
  • Measurement Method
    Name
    Factuality and Knowledge Robustness Testing 
    Applied Definition
    Testing whether system outputs remain accurate under, e.g., long-tail knowledge probing, temporal knowledge probing, numerical reasoning evaluation, or false-premise prompting.
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Red Teaming Evaluation Method
    References
    2refs
    • AI, NIST. "Artificial intelligence risk management framework: Generative artificial intelligence profile." NIST Trustworthy and Responsible AI Gaithersburg, MD, USA (2024)
    • Wang, Yuxia, Minghan Wang, Muhammad Arslan Manzoor, Fei Liu, Georgi Nenkov Georgiev, Rocktim Jyoti Das, and Preslav Nakov. "Factuality of large language models: A survey." In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 19519-19529. 2024.
    Tools
    3tools
    • microsoft/PyRIT
    • Language models, chat assistants, or coding agents
    • note-taking apps
  • Metric
    Name
    Fairness Through Awareness 
    Applied Definition
    A metric that tests whether similar individuals (x and x') are treated similarly by a model M, requiring a defined similarity metric (D) to compare individuals: d(M(x),M(x′)) ≤ D(x,x′).
    AI RMF Characteristic(s)
    Valid & ReliableFair
    AI Lifecycle Stage(s)
    Collect & Process DataBuildUseOperate & Monitor
    Primary TEVV Application
    Fairness Evaluation
    References
    1ref
    • Mehrabi, Ninareh, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. “A Survey on Bias and Fairness in Machine Learning.” ACM Computing Surveys (CSUR) 54 (6): 1–35. Link
    Tools
    1tool
  • Metric
    Name
    False Refusal Rate - Full (FRR-f) 
    Applied Definition
    Percentage of safe prompts that models completely refuse.
    AI RMF Characteristic(s)
    Safe
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Safety Evaluation
    References
    1ref
    • Röttger, Paul, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. 2024. “Xstest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.” Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5377–400. Link
  • Metric
    Name
    False Refusal Rate - Partial (FRR-p) 
    Applied Definition
    Percentage of safe prompts that models partially refuse (by not completing specific parts of the prompt).
    AI RMF Characteristic(s)
    Safe
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Safety Evaluation
    References
    1ref
    • Röttger, Paul, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. 2024. “Xstest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.” Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5377–400. Link
  • Metric
    Name
    Falsification Rate 
    Applied Definition
    Percentage of interactions where a model falsified information to achieve a goal.
    AI RMF Characteristic(s)
    Safe
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Safety Evaluation
    References
    1ref
    • Su, Zhe, Xuhui Zhou, Sanketh Rangreji, et al. 2025. “AI-LIEDAR: Examine the trade-off between utility and truthfulness in LLM agents.” Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 11867–94. Link
  • Metric
    Name
    Feature Importance Entropy 
    Applied Definition
    Entropy of the fraction of features used to explain each decision.
    AI RMF Characteristic(s)
    Explainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Interpretability & Explanation Evaluation
    References
    1ref
    • Umang Bhatt, Jose M. F. Moura, and Adrian Weller. Evaluating and Aggregating Feature-based Model Explanations. volume 3, pages 3016–3022, July 2020. doi: 10.24963/ijcai.2020/417. ISSN: 1045-0823. Link
  • Metric
    Name
    Feature Recall 
    Applied Definition
    Fraction of features used in a known model (gold standard) that are recovered by post-hoc explanation.
    AI RMF Characteristic(s)
    Valid & ReliableExplainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Interpretability & Explanation Evaluation
    References
    1ref
    • Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. ”Why Should I Trust you?” Explaining the Predictions of Any Classifier. In KDD 2016: Proceedings of the 22nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, August 2016. ACM. Link
    Tools
    5tools
    • scipy.org/
    • Derived from ground truth and explanation data
  • Measurement Method
    Name
    Field Pilots 
    Applied Definition
    Limited real-world deployments used to evaluate an AI system in operational settings, i.e., in-situ evaluation. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Valid & ReliableSafeSecure & ResilientAccountable & TransparentExplainable & InterpretablePrivacy-EnhancedFairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    Deploy
    Primary TEVV Application
    Human-Centered Evaluation
    References
    2refs
    • Glass, Robert L. "Pilot studies: What, why and how." Journal of Systems and Software 36, no. 1 (1997): 85-97
    • Commonwealth of Pennsylvania, Office of Administration, Lessons from Pennsylvania’s Generative AI Pilot with ChatGPT (March 2025).
  • Metric
    Name
    Fleiss' Kappa 
    Applied Definition
    An extension of Cohen's Kappa assessing the agreement reliability among a fixed number of three or more annotators when assigning categorical ratings: kappa = (P_bar - P^bar_e) / (1 - P^bar_e).
    AI RMF Characteristic(s)
    Valid & ReliableAccountable & Transparent
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Annotator Agreement Evaluation
    References
    1ref
    • Fleiss, Joseph L. 1971. “Measuring Nominal Scale Agreement Among Many Raters.” Psychological Bulletin 76 (5): 378. Link
  • Measurement Method
    Name
    Focus Groups and Interviews 
    Applied Definition
    Qualitative methods that gather detailed perspectives and experiences from participants about an AI system. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Valid & ReliableSafeSecure & ResilientAccountable & TransparentExplainable & InterpretablePrivacy-EnhancedFairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    Plan & DesignBuildUseOperate & Monitor
    Primary TEVV Application
    Human-Centered Evaluation
    References
    3refs
    • Ameller, David, Claudia Ayala, Jordi Cabot, and Xavier Franch. "How do software architects consider non-functional requirements: An exploratory study." In 2012 20th IEEE international requirements engineering conference (RE), pp. 41-50. IEEE, 2012
    • Beede, Emma, Elizabeth Baylor, Fred Hersch, Anna Iurchenko, Lauren Wilcox, Paisan Ruamviboonsuk, and Laura M. Vardoulakis. "A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy." In Proceedings of the 2020 CHI conference on human factors in computing systems, pp. 1-12. 2020
    • Chazette, Larissa, Jil Klünder, Merve Balci, and Kurt Schneider. "How can we develop explainable systems? insights from a literature review and an interview study." In Proceedings of the International Conference on Software and System Processes and International Conference on Global Software Engineering, pp. 1-12. 2022.
  • Measurement Method
    Name
    Garak 
    Applied Definition
    Security-focused scanning software package (Python) that runs many probes for vulnerabilities, unsafe outputs, and misuse-relevant failure modes.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Automated Language Model Red Teaming Toolkit
    References
    1ref
    • Derczynski, Leon, Erick Galinkin, Jeffrey Martin, Subho Majumdar, and Nanna Inie. "garak: A framework for security probing large language models." arXiv preprint arXiv:2406.11036 (2024).
    Tools
    1tool
  • Measurement Method
    Name
    GLUE / SuperGLUE 
    Applied Definition
    General language understanding evaluation (GLUE) benchmark with some hidden test set labels.
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    2refs
    • Wang, Alex, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. "GLUE: A multi-task benchmark and analysis platform for natural language understanding." In Proceedings of the 2018 EMNLP workshop BlackboxNLP: Analyzing and interpreting neural networks for NLP, pp. 353-355. 2018
    • Wang, Alex, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. "Superglue: A stickier benchmark for general-purpose language understanding systems." Advances in neural information processing systems 32 (2019).
  • Measurement Method
    Name
    GSM8K 
    Applied Definition
    Grade-school math word-problem benchmark. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Cobbe, Karl, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert et al. "Training verifiers to solve math word problems." arXiv preprint arXiv:2110.14168 (2021).
  • Metric
    Name
    Gwet's AC1 
    Applied Definition
    A paradox-resistant agreement coefficient for categorical data. It adjusts for chance agreement but avoids the "kappa paradox."
    AI RMF Characteristic(s)
    Valid & ReliableAccountable & Transparent
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Annotator Agreement Evaluation
    References
    1ref
    • Gwet, K. L. 2002. “Inter-Rater Reliability: Dependency on Trait Prevalence and Marginal Homogeneity.” Statistical Methods for Inter-Rater Reliability Assessment 2. Link
  • Measurement Method
    Name
    HarmBench 
    Applied Definition
    Benchmark for harmful request handling, refusal behavior, and automated red-teaming of safety failures. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    SafeSecure & Resilient
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Mazeika, Mantas, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee et al. "HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal." Proceedings of Machine Learning Research 235 (2024): 35181-35224.
  • Measurement Method
    Name
    HellaSwag 
    Applied Definition
    Benchmark for commonsense sentence completion. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Zellers, Rowan, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. "Hellaswag: Can a machine really finish your sentence?" In Proceedings of the 57th annual meeting of the association for computational linguistics, pp. 4791-4800. 2019.
  • Measurement Method
    Name
    HELM: Bias / Toxicity 
    Applied Definition
    HELM subset that evaluates biased model outputs across demographic or social dimensions and for toxicity or harm. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    FairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Bommasani, Rishi, Percy Liang, and Tony Lee. "Holistic Evaluation of Language Models." Annals of the New York Academy of Sciences 1525, no. 1 (July 2023): 140–146. Link
  • Metric
    Name
    Homomorphic Class 
    Applied Definition
    Categorization of encryption as fully, partially, or somewhat homomorphic based on operator limits; can be used to assess security risk.
    AI RMF Characteristic(s)
    Privacy-Enhanced
    AI Lifecycle Stage(s)
    Plan & DesignBuildUseDeploy
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    1ref
    • Abbas Acar, Hidayet Aksu, A. Selcuk Uluagac, and Mauro Conti. A Survey on Homomorphic Encryption Schemes: Theory and Implementation. ACM Computing Surveys, 51(4):79:1–79:35, July 2018. ISSN 0360-0300. doi: 10.1145/3214303. Link
    Tools
    4tools
    • Derived from encryption method
  • Measurement Method
    Name
    HumanEval 
    Applied Definition
    Code-generation benchmark. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Chen, Mark, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards et al. "Evaluating large language models trained on code." arXiv preprint arXiv:2107.03374 (2021).
  • Measurement Method
    Name
    Human Inspection Similarity 
    Applied Definition
    Qualitative metric checking if adversarial images resemble original classes. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Human-Centered Evaluation
    References
    1ref
    • Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Srndic, Pavel Laskov, Giorgio Giacinto, and Fabio Roli. Evasion attacks against machine learning at test time. In Joint European conference on machine learning and knowledge discovery in databases, pages 387–402. Springer, 2013.
  • Measurement Method
    Name
    Human Response Time 
    Applied Definition
    The duration a human takes to complete a simulation or prediction task. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Explainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Human-Centered Evaluation
    References
    2refs
    • Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Sam Gershman, and Finale Doshi-Velez. An Evaluation of the Human-Interpretability of Explanation. arXiv:1902.00006 [cs, stat], August 2019. arXiv: 1902.00006 Link
    • Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Samuel J. Gershman, and Finale Doshi-Velez. Human Evaluation of Models Built for Interpretability. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, 7(1):59–67, October 2019. Link
  • Metric
    Name
    Image Confidence Change 
    Applied Definition
    The difference in the model's confidence score for a specific class when noise (perturbation) is added to an image, moving from the original image to its perturbed version.
    AI RMF Characteristic(s)
    Valid & ReliableSecure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    1ref
    • Jennifer Fernick. Practical Attacks on Machine Learning Systems. NCC Group Research Whitepaper, July 2022. Link
  • Measurement Method
    Name
    ImageNet 
    Applied Definition
    Benchmark for large-scale image classification; also localization/detection in the ILSVRC challenge tasks. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Object Recognition Benchmark Evaluation
    References
    2refs
    • Deng, Jia, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. "Imagenet: A large-scale hierarchical image database." In 2009 IEEE conference on computer vision and pattern recognition, pp. 248-255. IEEE, 2009
    • Russakovsky, Olga, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang et al. "Imagenet large scale visual recognition challenge." International journal of computer vision 115, no. 3 (2015): 211-252.
    Tools
    1tool
  • Measurement Method
    Name
    Impact Assessment 
    Applied Definition
    A structured evaluation of an AI system’s likely impacts on people, organizations, and resources. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Valid & ReliableSafeSecure & ResilientAccountable & TransparentExplainable & InterpretablePrivacy-EnhancedFairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    Plan & Design
    Primary TEVV Application
    Human-Centered Evaluation
    References
    2refs
    • Reisman, Dillon, Jason Schultz, Kate Crawford, and Meredith Whittaker. "Algorithmic impact assessments: a practical framework for public agency." AI Now (2018)
    • Microsoft. Microsoft Responsible AI Impact Assessment Guide. June 2022.
  • Metric
    Name
    Infidelity 
    Applied Definition
    The expected difference between the dot product of the input perturbation to the explanation and the output perturbation.
    AI RMF Characteristic(s)
    Valid & ReliableExplainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Interpretability & Explanation Evaluation
    References
    1ref
    • Chih-Kuan Yeh, Cheng-Yu Hsieh, Arun Suggala, David I. Inouye, and Pradeep K. Ravikumar. On the (In)fidelity and Sensitivity of Explanations. pages 10967–10978, 2019. Link
  • Measurement Method
    Name
    Information Transfer Rate (ITR) 
    Applied Definition
    Mutual information between model and human predictions divided by response time. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Explainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Human-Centered Evaluation
    References
    1ref
    • Philipp Schmidt and Felix Biessmann. Quantifying Interpretability and Trust in Machine Learning Systems. arXiv:1901.08558 [cs, stat], January 2019. arXiv: 1901.08558. Link
  • Measurement Method
    Name
    Integrity Attacks 
    Applied Definition
    Testing whether a system can be manipulated to alter intended outputs or outcomes through, e.g., prompt injection, indirect prompt injection, repeated token attacks, data poisoning, encoded or obfuscated prompt attacks, or embedding/retrieval weakness.
    AI RMF Characteristic(s)
    Valid & ReliableSecure & ResilientAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Red Teaming Evaluation Method
    References
    6refs
    • Barreno, Marco, Blaine Nelson, Anthony D. Joseph, and J. Doug Tygar. "The security of machine learning." Machine learning 81, no. 2 (2010): 121-148
    • Liu, Yi, Gelei Deng, Yuekang Li, Kailong Wang, Zihao Wang, Xiaofeng Wang, Tianwei Zhang et al. "Prompt injection attack against LLM-integrated applications." arXiv preprint arXiv:2306.05499 (2023)
    • Greshake, Kai, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. "Not what you've signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection." In Proceedings of the 16th ACM workshop on artificial intelligence and security, pp. 79-90. 2023
    • Wan, Alexander, Eric Wallace, Sheng Shen, and Dan Klein. "Poisoning language models during instruction tuning." In International Conference on Machine Learning, pp. 35413-35425. PMLR, 2023
    • OWASP LLM01:2025 Prompt Injection, LLM04:2025 Data and Model Poisoning
    • MITRE ATLAS LLM Prompt Injection, AI Agent Context Poisoning, Poison Training Data, LLM Prompt Obfuscation, False RAG Entry Injection.
    Tools
    8tools
  • Metric
    Name
    Intraclass Correlation Coefficient (ICC) 
    Applied Definition
    A statistical metric used to evaluate the reliability and consistency of multiple annotators when the assigned labels or ratings are continuous, interval, or ratio data.
    AI RMF Characteristic(s)
    Valid & ReliableAccountable & Transparent
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Annotator Agreement Evaluation
    References
    1ref
    • Liljequist, David, Britt Elfving, and Kirsti Skavberg Roaldsen. 2019. “Intraclass correlation–A discussion and demonstration of basic features.” PloS One 14 (7): e0219854. Link
  • Metric
    Name
    Inversion Precision 
    Applied Definition
    Precision of predictions during a model inversion attack to learn sensitive features.
    AI RMF Characteristic(s)
    Privacy-Enhanced
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    1ref
    • Matt Fredrikson, Somesh Jha, and Thomas Ristenpart. Model Inversion Attacks that Exploit Confidence Information and Basic Countermeasures. In Proceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security, CCS ’15, pages 1322–1333, New York, NY, USA, October 2015. Association for Computing Machinery. ISBN 978-1-4503-3832-5. doi: 10.1145/2810103.2813677. Link
  • Metric
    Name
    Inversion Recall 
    Applied Definition
    Recall of predictions when measuring success of model inversion attacks.
    AI RMF Characteristic(s)
    Privacy-Enhanced
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    1ref
    • Matt Fredrikson, Somesh Jha, and Thomas Ristenpart. Model Inversion Attacks that Exploit Confidence Information and Basic Countermeasures. In Proceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security, CCS ’15, pages 1322–1333, New York, NY, USA, October 2015. Association for Computing Machinery. ISBN 978-1-4503-3832-5. doi: 10.1145/2810103.2813677. Link
  • Measurement Method
    Name
    Jailbreaking 
    Applied Definition
    Testing whether a system can be induced to violate safeguards through, e.g., prompt injection, multi-turn jailbreak prompting, poetic jailbreak prompting, benign framing, role-play, or task blending.
    AI RMF Characteristic(s)
    SafeSecure & Resilient
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Red Teaming Evaluation Method
    References
    4refs
    • Wei, Alexander, Nika Haghtalab, and Jacob Steinhardt. "Jailbroken: How does LLM safety training fail?" Advances in neural information processing systems 36 (2023): 80079-80110
    • Li, Nathaniel, Ziwen Han, Ian Steneker, Willow Primack, Riley Goodside, Hugh Zhang, Zifan Wang, Cristina Menghini, and Summer Yue. "LLM defenses are not robust to multi-turn human jailbreaks yet." arXiv preprint arXiv:2408.15221 (2024)
    • OWASP LLM01:2025 Prompt Injection
    • MITRE ATLAS LLM Prompt Injection, LLM Jailbreak.
    Tools
    7tools
    • Browser developer panes
    • Bash utilities
    • Language models, chat assistants, or coding agents
    • note-taking apps
  • Measurement Method
    Name
    JailbreakingLLMs 
    Applied Definition
    Python software, benchmarking methodology, attack set for evaluating susceptibility to jailbreak prompts.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Automated Language Model Red Teaming Toolkit
    References
    1ref
    • Chao, Patrick, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. "Jailbreaking black box large language models in twenty queries." In 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML). IEEE, 2025.
  • Metric
    Name
    Jailbreak Success Rate 
    Applied Definition
    Percentage of adversarial prompts that bypass built-in safety filters/guardrails.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Safety Evaluation
    References
    1ref
    • Wei, Alexander, Nika Haghtalab, and Jacob Steinhardt. 2023. “Jailbroken: How does LLM safety training fail?” Advances in Neural Information Processing Systems 36: 80079–110. Link
  • Metric
    Name
    Kendall's W 
    Applied Definition
    A normalization of the Friedman test statistic used to assess agreement among three or more annotators who are ranking a set of items (ordinal data scale).
    AI RMF Characteristic(s)
    Valid & ReliableAccountable & Transparent
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Annotator Agreement Evaluation
    References
    1ref
    • Kendall, Maurice G., and B. Babington Smith. 1939. “The Problem of m Rankings.” The Annals of Mathematical Statistics 10 (3): 275–87. Link
  • Metric
    Name
    Kernel Density Distance 
    Applied Definition
    Distance from an input to the kernel of the target class label.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    2refs
    • Reuben Feinman, Ryan R. Curtin, Saurabh Shintre, and Andrew B. Gardner. Detecting Adversarial Samples from Artifacts, November 2017. arXiv:1703.00410 [cs, stat] Link
    • Nicholas Carlini and David Wagner. Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security, AISec ’17, pages 3–14, New York, NY, USA, November 2017. Association for Computing Machinery. ISBN 978-1-4503-5202-4. doi: 10.1145/3128572.3140444. Link
  • Metric
    Name
    Krippendorff's Alpha 
    Applied Definition
    A versatile agreement coefficient for any number of annotators that handles missing data and adapts to nominal, ordinal, interval, and ratio scales: alpha = 1 - (D_o / D_e).
    AI RMF Characteristic(s)
    Valid & ReliableAccountable & Transparent
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Annotator Agreement Evaluation
    References
    1ref
    • Hayes, Andrew F., and Klaus Krippendorff. 2007. “Answering the Call for a Standard Reliability Measure for Coding Data.” Communication Methods and Measures 1 (1): 77–89. Link
  • Metric
    Name
    L₀ Norm 
    Applied Definition
    Number of features/pixels modified in an adversarial attack.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    1ref
    • Nicholas Carlini and David Wagner. Towards Evaluating the Robustness of Neural Networks, March 2017. arXiv:1608.04644 [cs]. Link
  • Metric
    Name
    L₂ Norm 
    Applied Definition
    Euclidean distance between original and adversarial input vectors.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    2refs
    • Nicholas Carlini and David Wagner. Towards Evaluating the Robustness of Neural Networks, March 2017. arXiv:1608.04644 [cs] Link
    • Pin-Yu Chen, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh. ZOO: Zeroth Order Optimization Based Black-box Attacks to Deep Neural Networks without Training Substitute Models. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security, AISec ’17, pages 15–26, New York, NY, USA, November 2017. Association for Computing Machinery. ISBN 978-1-4503-5202-4. doi: 10.1145/3128572.3140448. Link
  • Metric
    Name
    Linear Regression Coefficients 
    Applied Definition
    The estimated weights (beta_i) representing the expected change in the dependent variable per unit change in a predictor: Y_hat = beta_0 + beta_1*X_1 + ... + beta_k*X_k. Typically used after an unconditioned metric (e.g., Cohen's d) to assess whether a demographic indicator x_j is a statistically significant predictor of Y when conditioning on additional relevant variables.
    AI RMF Characteristic(s)
    Valid & ReliableFair
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Fairness Evaluation
    References
    1ref
    • Hastie, Trevor, Robert Tibshirani, and Jerome Friedman. 2009. The Elements of Statistical Learning: Data Mining, Inference, and Prediction. 2nd ed. Springer Series in Statistics. Springer. Link
  • Measurement Method
    Name
    LiveBench 
    Applied Definition
    Continuously refreshed benchmark designed to reduce contamination and test current general-purpose knowledge.
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • White, Colin, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain, Ravid Shwartz-Ziv et al. "LiveBench: A Challenging, Contamination-Limited LLM Benchmark." In The Thirteenth International Conference on Learning Representations.
  • Measurement Method
    Name
    llmprivacy 
    Applied Definition
    Software and prompting approaches for assessing privacy leakage or privacy risks in large language model outputs.
    AI RMF Characteristic(s)
    Secure & ResilientPrivacy-Enhanced
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Automated Language Model Red Teaming Toolkit
    References
    1ref
    • Staab, Robin, Mark Vero, Mislav Balunović, and Martin Vechev. "Beyond Memorization: Violating Privacy Via Inference with Large Language Models." In The Twelfth International Conference on Learning Representations. OpenReview, 2024.
  • Metric
    Name
    L∞ Norm 
    Applied Definition
    Maximum absolute change to any single input feature.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    1ref
    • Nicholas Carlini and David Wagner. Towards Evaluating the Robustness of Neural Networks, March 2017. arXiv:1608.04644 [cs]. Link
  • Metric
    Name
    Local Intrinsic Dimensionality (LID) 
    Applied Definition
    Metric characterizing the area/subspace around an input.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeployOperate & Monitor
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    1ref
    • Xingjun Ma, Bo Li, Yisen Wang, Sarah M. Erfani, Sudanthi Wijewickrema, Grant Schoenebeck, Dawn Song, Michael E. Houle, and James Bailey. Characterizing Adversarial Subspaces Using Local Intrinsic Dimensionality, March 2018. arXiv:1801.02613 [cs]. Link
  • Metric
    Name
    Logistic Regression Coefficients 
    Applied Definition
    The weights (beta_i) representing the change in the log-odds of the outcome per unit change in the predictor: ln(p/(1-p)) = beta_0 + beta_1*X_1. Evaluates directional risk factors. Typically used after an unconditioned metric (e.g., demographic parity) to assess whether a demographic indicator x_j is a statistically significant predictor of Y when conditioning on additional relevant variables.
    AI RMF Characteristic(s)
    Valid & ReliableFair
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Fairness Evaluation
    References
    1ref
    • Hastie, Trevor, Robert Tibshirani, and Jerome Friedman. 2009. The Elements of Statistical Learning: Data Mining, Inference, and Prediction. 2nd ed. Springer Series in Statistics. Springer. Link
  • Measurement Method
    Name
    Malicious Code Elicitation 
    Applied Definition
    Testing whether a model or agent can be induced to produce malicious, exploitative, or covertly harmful code through, e.g., improper output artifact elicitation, malware generation, hidden payloads, unsafe scripting, or backdoored implementations.
    AI RMF Characteristic(s)
    SafeSecure & Resilient
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Red Teaming Evaluation Method
    References
    4refs
    • Khoury, Raphaël, Anderson R. Avila, Jacob Brunelle, and Baba Mamadou Camara. "How secure is code generated by ChatGPT?" arXiv preprint arXiv:2304.09655 (2023)
    • Li, Haoyang, Huan Gao, Zhiyuan Zhao, Zhiyu Lin, Junyu Gao, and Xuelong Li. "LLMs caught in the crossfire: Malware requests and jailbreak challenges." In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 27833-27848. 2025
    • Ouyang, Sheng, Yihao Qin, Bo Lin, Liqian Chen, Xiaoguang Mao, and Shangwen Wang. "Smoke and mirrors: Jailbreaking LLM-based code generation via implicit malicious prompts." arXiv preprint arXiv:2503.17953 (2025)
    • OWASP LLM05:2025 Improper Output Handling. MITRE ATLAS Generate Malicious Commands.
    Tools
    9tools
  • Metric
    Name
    Maximum Sensitivity 
    Applied Definition
    Peak change in explanation distance within a local neighborhood of inputs.
    AI RMF Characteristic(s)
    Valid & ReliableExplainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeployOperate & Monitor
    Primary TEVV Application
    Interpretability & Explanation Evaluation
    References
    1ref
    • Umang Bhatt, José M. F. Moura, and Adrian Weller. Evaluating and Aggregating Feature-based Model Explanations. volume 3, pages 3016–3022, July 2020. doi: 10.24963/ijcai.2020/417. ISSN: 1045-0823. Link
  • Metric
    Name
    Mean Absolute Pixel Error 
    Applied Definition
    Average error between decoded images and actual images in memorization attacks.
    AI RMF Characteristic(s)
    Privacy-Enhanced
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    1ref
    • Congzheng Song, Thomas Ristenpart, and Vitaly Shmatikov. Machine Learning Models that Remember Too Much. In Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, CCS ’17, pages 587– 601, New York, NY, USA, October 2017. Association for Computing Machinery. ISBN 978-1-4503-4946-8. doi: 10.1145/3133956.3134077. Link
  • Metric
    Name
    Mean Recall 
    Applied Definition
    The average fraction of "gold features" (the features actually known to be used by a model) that an explanation algorithm successfully identifies across a set of test instances.
    AI RMF Characteristic(s)
    Valid & ReliableExplainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Interpretability & Explanation Evaluation
    References
    1ref
    • Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. ”Why Should I Trust you?” Explaining the Predictions of Any Classifier. In KDD 2016: Proceedings of the 22nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, August 2016. ACM. Link
    Tools
    5tools
    • scipy.org/
    • Derived from ground truth and explanation data
  • Metric
    Name
    Median Absolute Deviation (MAD) 
    Applied Definition
    Distance metric measuring the spread between original and altered counterfactual inputs.
    AI RMF Characteristic(s)
    Explainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Interpretability & Explanation Evaluation
    References
    1ref
    • S. Wachter, B. D. M. Mittelstadt, and C. Russell. Counterfactual explanations without opening the black box: automated decisions and the GDPR. Harvard Journal of Law and Technology, 31(2), 2018. ISSN 0897-3393. Link
  • Metric
    Name
    Median Squared L_2 Distance 
    Applied Definition
    A metric quantifying attack quality by measuring the median amount of perturbation (expressed as L_2 norms) required to cause a model to misclassify a set of test instances.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    1ref
    • Wieland Brendel, Jonas Rauber, and Matthias Bethge. Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning Models, February 2018. arXiv:1712.04248 [cs, stat]. Link
  • Measurement Method
    Name
    mimir 
    Applied Definition
    Python-based software for measuring memorization in LLMs.
    AI RMF Characteristic(s)
    Secure & ResilientPrivacy-Enhanced
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Automated Language Model Red Teaming Toolkit
    References
    1ref
    • Duan, Michael, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettlemoyer, Yulia Tsvetkov, Yejin Choi, David Evans, and Hannaneh Hajishirzi. "Do Membership Inference Attacks Work on Large Language Models?" In First Conference on Language Modeling.
  • Measurement Method
    Name
    MLCommons: AILuminate 
    Applied Definition
    Evaluation for safety-related model risks and performance under standardized testing setups; controlled test data.
    AI RMF Characteristic(s)
    Safe
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Ghosh, Shaona, Heather Frase, Adina Williams, Sarah Luger, Paul Röttger, Fazl Barez, Sean McGregor et al. "AILuminate: Introducing v1.0 of the AI risk and reliability benchmark from MLCommons." arXiv preprint arXiv:2503.05731 (2025).
  • Measurement Method
    Name
    MMLU 
    Applied Definition
    Broad multitask benchmark. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Hendrycks, Dan, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. "Measuring Massive Multitask Language Understanding." In International Conference on Learning Representations.
    Tools
    1tool
  • Measurement Method
    Name
    MS COCO 
    Applied Definition
    Object recognition benchmark with various tasks, including captioning. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Object Recognition Benchmark Evaluation
    References
    1ref
    • Lin, Tsung-Yi, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. "Microsoft COCO: Common objects in context." In European conference on computer vision, pp. 740-755. Cham: Springer International Publishing, 2014.
  • Measurement Method
    Name
    MT-bench 
    Applied Definition
    Multi-turn chat benchmark. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Zheng, Lianmin, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin et al. "Judging LLM-as-a-judge with MT-Bench and chatbot arena." Advances in neural information processing systems 36 (2023): 46595-46623.
    Tools
    1tool
  • Measurement Method
    Name
    NIST FRTE/FATE 
    Applied Definition
    NIST evaluates systems on facial recognition tasks using standardized, sequestered test data not available to developers.
    AI RMF Characteristic(s)
    Valid & ReliableFairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    Operate & Monitor
    Primary TEVV Application
    Facial Recognition Benchmark Evaluation
    References
    3refs
    • Ngan, Mei, Patrick Grother, Kayee Hanaoka, and Jason Kuo. "NISTIR 8292 DRAFT SUPPLEMENT, Face Analysis Technology Evaluation (FATE) Part 4: MORPH-Performance of Automated Face Morph Detection." US Department of Commerce, National Institute of Standards and Technology (2026)
    • Grother, Patrick, Austin Hom, Mei Ngan, Kayee Hanaoka. "Draft for Public Comment, Ongoing Face Recognition Vendor Test (FRVT) Part 5: Face Image Quality Assessment." US Department of Commerce, National Institute of Standards and Technology (2022)
    • NIST, "Face Analysis Technology Evaluation (FATE) Quality Assessment - Specific Image Defect Detection, API and Concept Document." US Department of Commerce, National Institute of Standards and Technology (2024).
  • Measurement Method
    Name
    Normalized Difficulty Rating 
    Applied Definition
    Subjective difficulty normalized by individual subject means across tasks. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Explainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Human-Centered Evaluation
    References
    2refs
    • Isaac Lage, Emily Chen, Jeffrey He,Menaka Narayanan, Been Kim, Sam Gershman, and Finale Doshi-Velez. An Evaluation of the Human-Interpretability of Explanation. arXiv:1902.00006 [cs, stat], August 2019. arXiv: 1902.00006 Link
    • Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Samuel J. Gershman, and Finale Doshi-Velez. Human Evaluation of Models Built for Interpretability. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, 7(1):59–67, October 2019. Link
  • Measurement Method
    Name
    Normalized Trust 
    Applied Definition
    Psychometric measurement of a user's confidence in an AI system's output, collected via subjective rating scales (such as a Likert scale) and then normalized at the individual subject level to minimize variance caused by personal rating behaviors. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Accountable & Transparent
    AI Lifecycle Stage(s)
    BuildUseDeployOperate & Monitor
    Primary TEVV Application
    Human-Centered Evaluation
    References
    1ref
    • Jianlong Zhou, Zhidong Li, Huaiwen Hu, Kun Yu, Fang Chen, Zelin Li, and Yang Wang. Effects of Influence on User Trust in Predictive Decision Making. In Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems, CHI EA ’19, pages 1–6, New York, NY, USA, May 2019. Association for Computing Machinery. ISBN 978-1-4503-5971-9. doi: 10.1145/3290607.3312962. Link
  • Metric
    Name
    Number of Coefficients 
    Applied Definition
    Total count of feature coefficients in a logistic regression model.
    AI RMF Characteristic(s)
    Explainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Interpretability & Explanation Evaluation
    References
    1ref
    • Forough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan, and Hanna Wallach. Manipulating and Measuring Model Interpretability. arXiv:1802.07810 [cs], November 2019. arXiv: 1802.07810. Link
  • Metric
    Name
    Number of Rows Shared 
    Applied Definition
    Count of data rows shared during federated training.
    AI RMF Characteristic(s)
    Privacy-Enhanced
    AI Lifecycle Stage(s)
    BuildUse
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    1ref
    • Mohammed Aledhari, Rehma Razzak, Reza M. Parizi, and Fahad Saeed. Federated Learning: A Survey on Enabling Technologies, Protocols, and Applications. IEEE Access, 8:140699–140725, 2020. ISSN 2169-3536. doi: 10.1109/ACCESS.2020.3013541. Conference Name: IEEE Access. Link
    Tools
    3tools
  • Measurement Method
    Name
    OpenBookQA 
    Applied Definition
    Benchmark for open-book science question answering. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Mihaylov, Todor, Peter Clark, Tushar Khot, and Ashish Sabharwal. "Can a suit of armor conduct electricity? A new dataset for open book question answering." In Proceedings of the 2018 conference on empirical methods in natural language processing, pp. 2381-2391. 2018.
  • Metric
    Name
    Paraphrased Plagiarism Rate 
    Applied Definition
    Percentage of documents that LLMs generate with paraphrasing.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    Collect & Process DataBuildUseOperate & Monitor
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    3refs
    • Lee, Jooyoung, Thai Le, Jinghui Chen, and Dongwon Lee. 2023. “Do Language Models Plagiarize?” Proceedings of the ACM Web Conference 2023, 3637–47. Link
    • Chen, Tong, Faeze Brahman, Jiacheng Liu, et al. 2025. ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-Training Data. Link
    • Lee, Jooyoung, Toshini Agrawal, Adaku Uchendu, Thai Le, Jinghui Chen, and Dongwon Lee. 2025. “PlagBench: Exploring the Duality of Large Language Models in Plagiarism Generation and Detection.” In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), edited by Luis Chiruzzo, Alan Ritter, and Lu Wang. Association for Computational Linguistics. Link
  • Metric
    Name
    Partial Refusal Rate 
    Applied Definition
    Percentage of prompts that models partially refuse (by not completing specific parts of the prompt).
    AI RMF Characteristic(s)
    Safe
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Safety Evaluation
    References
    1ref
    • Röttger, Paul, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. 2024. “Xstest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.” Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5377–400. Link
  • Measurement Method
    Name
    PASCAL VOC 
    Applied Definition
    Object recognition benchmark with publicly distributed standard datasets and annotations. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Object Recognition Benchmark Evaluation
    References
    1ref
    • Everingham, Mark, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. "The pascal visual object classes (voc) challenge." International journal of computer vision 88, no. 2 (2010): 303-338.
  • Metric
    Name
    Percent Agreement 
    Applied Definition
    The most straightforward agreement metric, calculated as the percentage of total items for which all annotators assigned the exact same label without correcting for chance.
    AI RMF Characteristic(s)
    Accountable & Transparent
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Annotator Agreement Evaluation
    References
    1ref
    • Artstein, Ron, and Massimo Poesio. 2008. “Survey Article: Inter-Coder Agreement for Computational Linguistics.” Computational Linguistics 34 (4): 555–96. Link
  • Metric
    Name
    Per-Example Rate 
    Applied Definition
    Success rate reported for each individual example rather than as an average.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    2refs
    • Anish Athalye, Nicholas Carlini, and David Wagner. Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples, July 2018. arXiv:1802.00420 [cs] Link
    • Nicholas Carlini, Anish Athalye, Nicolas Papernot, Wieland Brendel, Jonas Rauber, Dimitris Tsipras, Ian Goodfellow, Aleksander Madry, and Alexey Kurakin. On Evaluating Adversarial Robustness. arXiv:1902.06705 [cs, stat], February 2019. arXiv: 1902.06705. Link
  • Measurement Method
    Name
    PMLB 
    Applied Definition
    A large, curated repository of benchmark datasets for evaluating supervised machine learning algorithms. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Tabular Data Benchmark Evaluation
    References
    1ref
    • Olson, Randal S., William La Cava, Patryk Orzechowski, Ryan J. Urbanowicz, and Jason H. Moore. "PMLB: a large benchmark suite for machine learning evaluation and comparison." BioData mining 10, no. 1 (2017): 36.
  • Metric
    Name
    Poisoning Classification Accuracy 
    Applied Definition
    Fraction of models correctly identified as poisoned or clean.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    1ref
    • Peter Bajcsy and Michael Majurski. Baseline Pruning-Based Approach to Trojan Detection in Neural Networks. In ICLR 2021 Workshop on Security and Safety in Machine Learning Systems, page 5, 2021. Link
  • Measurement Method
    Name
    Post-deployment Feedback 
    Applied Definition
    Analysis of feedback after release to identify positive and negative real-world impacts, e.g., ratings, complaints, and incidents analysis. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Valid & ReliableSafeSecure & ResilientAccountable & TransparentExplainable & InterpretablePrivacy-EnhancedFairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    Operate & Monitor
    Primary TEVV Application
    Human-Centered Evaluation
    References
    2refs
    • Dai, Jessica, Inioluwa Deborah Raji, Benjamin Recht, and Irene Y. Chen. "Aggregated individual reporting for post-deployment evaluation." arXiv preprint arXiv:2506.18133 (2025)
    • McGregor, Sean. "Preventing repeated real world AI failures by cataloging incidents: The AI incident database." In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 17, pp. 15458-15463. 2021.
  • Measurement Method
    Name
    PyRIT 
    Applied Definition
    Framework for automated and human-led AI red teaming.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Automated Language Model Red Teaming Toolkit
    References
    1ref
    • Munoz, Gary D. Lopez, Amanda J. Minnich, Roman Lutz, Richard Lundeen, Raja Sekhar Rao Dheekonda, Nina Chikanov, Bolor-Erdene Jagdagdorj et al. "Pyrit: A framework for security risk identification and red teaming in generative AI systems." arXiv preprint arXiv:2410.02828 (2024).
    Tools
    1tool
  • Metric
    Name
    R-ASR 
    Applied Definition
    Random Access Success Rate checking how many random samples are adversarial.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    1ref
    • Roland S. Zimmermann, Wieland Brendel, Florian Tramer, and Nicholas Carlini. Increasing Confidence in Adversarial Robustness Evaluations, June 2022. arXiv:2206.13991 [cs]. Link
  • Metric
    Name
    Refusal Rate 
    Applied Definition
    Percentage of prompts that models refuse.
    AI RMF Characteristic(s)
    Safe
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Safety Evaluation
    References
    1ref
    • Röttger, Paul, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. 2024. “Xstest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.” Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5377–400. Link
  • Metric
    Name
    Relative Estimation Error 
    Applied Definition
    Estimation error of hyperparameters when evaluating adversarial attacks.
    AI RMF Characteristic(s)
    Privacy-Enhanced
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    1ref
    • Binghui Wang and Neil Zhenqiang Gong. Stealing Hyperparameters in Machine Learning. In 2018 IEEE Symposium on Security and Privacy (SP), pages 36–52, May 2018. doi: 10.1109/SP.2018.00038. ISSN: 2375-1207.
  • Metric
    Name
    Response Stance 
    Applied Definition
    Numerical or categorical stance that model takes on an issue.
    AI RMF Characteristic(s)
    Valid & ReliableFair
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Safety Evaluation
    References
    1ref
    • Röttger, Paul, Musashi Hinck, Valentin Hofmann, et al. 2026. “IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance.” Transactions of the Association for Computational Linguistics 14: 318–40. Link
  • Metric
    Name
    RMSE (Private Regression) 
    Applied Definition
    Measure of the root mean square error of an AI system's predictions on a regression task, used to characterize the performance loss incurred when privacy-preserving mechanisms (e.g., differential privacy with varying epsilon values) are applied.
    AI RMF Characteristic(s)
    Valid & ReliablePrivacy-Enhanced
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    1ref
    • Harsha Nori, Rich Caruana, Zhiqi Bu, Judy Hanwen Shen, and Janardhan Kulkarni. Accuracy, Interpretability, and Differential Privacy via Explainable Boosting. arXiv:2106.09680 [cs], June 2021. arXiv: 2106.09680. Link
  • Measurement Method
    Name
    Robustness Certificate 
    Applied Definition
    A theoretical proof that identifies the minimum distance (or similar), represented by the value p_{cert}, from the decision boundary to a given input.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    2refs
    • Sungyoon Lee, Woojin Lee, Jinseong Park, and Jaewook Lee. Towards Better Understanding of Training Certifiably Robust Models against Adversarial Examples. In Advances in Neural Information Processing Systems, volume 34, pages 953–964. Curran Associates, Inc., 2021. Link
    • Sahil Singla and Soheil Feizi. Second-Order Provable Defenses against Adversarial Attacks, June 2020. arXiv:2006.00731 [cs, stat]. Link
  • Metric
    Name
    Rule Length 
    Applied Definition
    Number of clauses contained within each individual rule.
    AI RMF Characteristic(s)
    Explainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Interpretability & Explanation Evaluation
    References
    1ref
    • Himabindu Lakkaraju, Stephen H. Bach, and Jure Leskovec. Interpretable Decision Sets: A Joint Framework for Description and Prediction. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’16, pages 1675–1684, New York, NY, USA, 2016. Association for Computing Machinery. ISBN 978-1-4503-4232-2. doi: 10.1145/2939672.2939874. event-place: San Francisco, California, USA. Link
    Tools
    3tools
  • Metric
    Name
    Runtime Overhead 
    Applied Definition
    Computational delay added by a defense mechanism.
    AI RMF Characteristic(s)
    Valid & ReliableSecure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    1ref
    • Yannik Potdevin, Dirk Nowotka, and Vijay Ganesh. An Empirical Investigation of Randomized Defenses against Adversarial Attacks, September 2019. arXiv:1909.05580 [cs, stat]. Link
  • Measurement Method
    Name
    Sanity Check Success 
    Applied Definition
    Human or automated verification of whether features highlighted by an explanation are logically relevant (e.g., foreground vs. background). Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Accountable & TransparentExplainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeployOperate & Monitor
    Primary TEVV Application
    Interpretability & Explanation Evaluation
    References
    3refs
    • Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. ”Why Should I Trust you?” Explaining the Predictions of Any Classifier. In KDD 2016: Proceedings of the 22nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, August 2016. ACM. Link
    • Sina Mohseni, Niloofar Zarei, and Eric D. Ragan. A Multidisciplinary Survey and Framework for Design and Evaluation of Explainable AI Systems. arXiv:1811.11839 [cs], August 2020. arXiv: 1811.11839 Link
    • Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. Sanity Checks for Saliency Maps. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31, pages 9505–9515. Curran Associates, Inc., 2018. Link
  • Metric
    Name
    Scott's Pi 
    Applied Definition
    An agreement metric for two annotators and nominal data. Similar to Cohen's Kappa, but it assumes both annotators share the same baseline probability distribution for the categories.
    AI RMF Characteristic(s)
    Valid & ReliableAccountable & Transparent
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Annotator Agreement Evaluation
    References
    1ref
    • Scott, William A. 1955. “Reliability of Content Analysis: The Case of Nominal Scale Coding.” Public Opinion Quarterly, 321–25. Link
  • Metric
    Name
    Shared Data Size (Bytes) 
    Applied Definition
    Size of data communicated between machines in federated learning.
    AI RMF Characteristic(s)
    Privacy-Enhanced
    AI Lifecycle Stage(s)
    BuildUse
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    1ref
    • Runhua Xu, Nathalie Baracaldo, Yi Zhou, Ali Anwar, James Joshi, and Heiko Ludwig. FedV: Privacy-Preserving Federated Learning over Vertically Partitioned Data. In Proceedings of the 14th ACM Workshop on Artificial Intelligence and Security, AISec ’21, pages 181–192, New York, NY, USA, November 2021. Association for Computing Machinery. ISBN 978-1-4503-8657-9. doi: 10.1145/3474369.3486872. Link
    Tools
    3tools
  • Measurement Method
    Name
    Significant Regression Coefficients 
    Applied Definition
    Boolean metric (p < 0.05) indicating if feature weights significantly impact subjective difficulty. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Explainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Human-Centered Evaluation
    References
    2refs
    • Isaac Lage, Emily Chen, Jeffrey He,Menaka Narayanan, Been Kim, Sam Gershman, and Finale Doshi-Velez. An Evaluation of the Human-Interpretability of Explanation. arXiv:1902.00006 [cs, stat], August 2019. arXiv: 1902.00006 Link
    • Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Samuel J. Gershman, and Finale Doshi-Velez. Human Evaluation of Models Built for Interpretability. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, 7(1):59–67, October 2019. Link
  • Metric
    Name
    Standard Deviation Distance 
    Applied Definition
    Variability of the alteration distance across adversarial samples.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    1ref
    • Nicholas Carlini and David Wagner. Towards Evaluating the Robustness of Neural Networks, March 2017. arXiv:1608.04644 [cs]. Link
  • Metric
    Name
    Structural Similarity Index (SSIM) 
    Applied Definition
    Measures local patch luminance, contrast, and structure similarities between images to assess image-based explanations.
    AI RMF Characteristic(s)
    Valid & ReliableExplainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Interpretability & Explanation Evaluation
    References
    3refs
    • Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. Sanity Checks for Saliency Maps. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31, pages 9505–9515. Curran Associates, Inc., 2018. Link
    • Leon Sixt, Maximilian Granz, and Tim Landgraf. When Explanations Lie: Why Many Modified BP Attributions Fail. arXiv, December 2019. Link
    • Zhou Wang and Alan C. Bovik. Mean squared error: Love it or leave it? A new look at Signal Fidelity Measures. IEEE Signal Processing Magazine, 26(1):98–117, January 2009. ISSN 1558-0792. doi: 10.1109/MSP.2008.930649. Conference Name: IEEE Signal Processing Magazine.
  • Measurement Method
    Name
    Structured Human-subject Experiments 
    Applied Definition
    Controlled studies with human participants used to measure how people interact with or are affected by an AI system. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Valid & ReliableSafeSecure & ResilientAccountable & TransparentExplainable & InterpretablePrivacy-EnhancedFairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Human-Centered Evaluation
    References
    2refs
    • Schwartz, Reva, Gabriella Waters, Razvan Amironesei, Craig Greenberg, Jon Fiscus, Patrick Hall, Anya Jones et al. "The Assessing Risks and Impacts of AI (ARIA) Program Evaluation Design Document." (2024)
    • Amironesei, Razvan, Afzal Godil, Craig Greenberg, Kristen Greene, Theodore Jensen, Patrick Hall et. al. "Assessing Risks and Impacts of AI (ARIA) ARIA 0.1: Pilot Evaluation Report" NIST AI 700-2, Gaithersburg, MD, USA (2025).
  • Measurement Method
    Name
    Subjective Difficulty 
    Applied Definition
    Numeric rating of user-perceived difficulty in simulating a system. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Explainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Human-Centered Evaluation
    References
    2refs
    • Isaac Lage, Emily Chen, Jeffrey He,Menaka Narayanan, Been Kim, Sam Gershman, and Finale Doshi-Velez. An Evaluation of the Human-Interpretability of Explanation. arXiv:1902.00006 [cs, stat], August 2019. arXiv: 1902.00006 Link
    • Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Samuel J. Gershman, and Finale Doshi-Velez. Human Evaluation of Models Built for Interpretability. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, 7(1):59–67, October 2019. Link
  • Measurement Method
    Name
    SWE-bench 
    Applied Definition
    Software engineering benchmark; variants offer partially closed or closed assessment data.
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Jimenez, Carlos E., John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R. Narasimhan. "SWE-bench: Can Language Models Resolve Real-world Github Issues?." In The Twelfth International Conference on Learning Representations.
  • Measurement Method
    Name
    System Causability Scale (SCS) 
    Applied Definition
    Questionnaire providing a numeric rating for the quality of causal explanations. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Explainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Human-Centered Evaluation
    References
    1ref
    • Andreas Holzinger, Andre Carrington, and Heimo Muller. Measuring the Quality of Explanations: The System Causability Scale (SCS). KI - Kunstliche Intelligenz, 34(2):193–198, June 2020. ISSN 1610-1987. doi: 10.1007/s13218-020-00636-z. Link
  • Measurement Method
    Name
    TabArena 
    Applied Definition
    Maintained tabular data benchmark with public datasets and analysis artifacts. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Tabular Data Benchmark Evaluation
    References
    1ref
    • Erickson, Nick, Lennart Purucker, Andrej Tschalzev, David Holzmüller, Prateek Mutalik Desai, David Salinas, and Frank Hutter. "TabArena: A Living Benchmark for Machine Learning on Tabular Data." In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  • Measurement Method
    Name
    TabReD 
    Applied Definition
    Tabular data benchmark composed of larger datasets with some realistic characteristics. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Tabular Data Benchmark Evaluation
    References
    1ref
    • Rubachev, Ivan, Nikolay Kartashev, Yury Gorishniy, and Artem Babenko. "TabReD: Analyzing Pitfalls and Filling the Gaps in Tabular Deep Learning Benchmarks." In The Thirteenth International Conference on Learning Representations.
  • Measurement Method
    Name
    TAP 
    Applied Definition
    Query-efficient jailbreak method used to evaluate black-box model vulnerability to adversarial prompting.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Automated Language Model Red Teaming Toolkit
    References
    1ref
    • Mehrotra, Anay, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi. "Tree of attacks: Jailbreaking black-box llms automatically." Advances in Neural Information Processing Systems 37 (2024): 61065-61105.
    Tools
    1tool
  • Metric
    Name
    Target Class Misclassification Rate 
    Applied Definition
    Fraction of examples misclassified as a specifically desired target class.
    AI RMF Characteristic(s)
    Valid & ReliableSecure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    4refs
    • Noam Yefet, Uri Alon, and Eran Yahav. Adversarial examples for models of code. Proceedings of the ACM on Programming Languages, 4(OOPSLA):162:1–162:30, November 2020. doi: 10.1145/3428230. Link
    • Pin-Yu Chen, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh. ZOO: Zeroth Order Optimization Based Black-box Attacks to Deep Neural Networks without Training Substitute Models. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security, AISec ’17, pages 15–26, New York, NY, USA, November 2017. Association for Computing Machinery. ISBN 978-1-4503-5202-4. doi:10.1145/3128572.3140448. Link
    • Jiliang Zhang and Chen Li. Adversarial Examples: Opportunities and Challenges. IEEE Transactions on Neural Networks and Learning Systems, 31(7):2578–2593, July 2020. ISSN 2162-2388. doi: 10.1109/TNNLS.2019.2933524. Conference Name: IEEE Transactions on Neural Networks and Learning Systems
    • Nicholas Carlini and David Wagner. Towards Evaluating the Robustness of Neural Networks, March 2017. arXiv:1608.04644 [cs]. Link
  • Metric
    Name
    Test Fairness (Well-Calibration) 
    Applied Definition
    For probabilities P, labels Y, group A, predicted probability score S, people in all groups must have an equal probability of correctly belonging to the positive class (P(Y=1|S=s,A=0)=P(Y=1|S=s,A=1).
    AI RMF Characteristic(s)
    Valid & ReliableFair
    AI Lifecycle Stage(s)
    Collect & Process DataBuildUseOperate & Monitor
    Primary TEVV Application
    Fairness Evaluation
    References
    1ref
    • Mehrabi, Ninareh, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. “A Survey on Bias and Fairness in Machine Learning.” ACM Computing Surveys (CSUR) 54 (6): 1–35. Link
  • Metric
    Name
    Toxicity Score 
    Applied Definition
    Level of toxicity in the generated content (with toxicity defined loosely as hate-speech, profanity, etc.).
    AI RMF Characteristic(s)
    Safe
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Safety Evaluation
    References
    1ref
    • Gehman, Samuel, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020. “Realtoxicityprompts: Evaluating Neural Toxic Degeneration in Language Models.” Findings of the Association for Computational Linguistics: EMNLP 2020, 3356–69. Link
  • Measurement Method
    Name
    ToxiGen 
    Applied Definition
    Benchmark for identifying toxic or hate-related content. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    FairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Hartvigsen, Thomas, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. "Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection." In Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: Long papers), pp. 3309-3326. 2022.
  • Metric
    Name
    Transferability Fraction 
    Applied Definition
    Fraction of examples generated for Model i that misclassify Model j.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    BuildUseDeploy
    Primary TEVV Application
    Adversarial Robustness Evaluation
    References
    2refs
    • Nicolas Papernot, Patrick McDaniel, and Ian Goodfellow. Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples, May 2016. arXiv:1605.07277 [cs] Link
    • Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks, February 2014. arXiv:1312.6199 [cs]. Link
  • Measurement Method
    Name
    Transformed Task Accuracy 
    Applied Definition
    Accuracy on a transformed task in which explanations are provided as part of the model output to support user decision-making. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Valid & ReliableExplainable & Interpretable
    AI Lifecycle Stage(s)
    BuildUseDeployOperate & Monitor
    Primary TEVV Application
    Human-Centered Evaluation
    References
    2refs
    • Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. ”Why Should I Trust you?” Explaining the Predictions of Any Classifier. In KDD 2016: Proceedings of the 22nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, August 2016. ACM. Link
    • Sina Mohseni, Niloofar Zarei, and Eric D. Ragan. A Multidisciplinary Survey and Framework for Design and Evaluation of Explainable AI Systems. arXiv:1811.11839 [cs], August 2020. arXiv: 1811.11839. Link
  • Metric
    Name
    Treatment Equality 
    Applied Definition
    Achieved when the ratio of false negatives to false positives is the same for both protected and unprotected group categories.
    AI RMF Characteristic(s)
    Valid & ReliableFair
    AI Lifecycle Stage(s)
    Collect & Process DataBuildUseOperate & Monitor
    Primary TEVV Application
    Fairness Evaluation
    References
    1ref
    • Mehrabi, Ninareh, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. “A Survey on Bias and Fairness in Machine Learning.” ACM Computing Surveys (CSUR) 54 (6): 1–35. Link
  • Metric
    Name
    Truthfulness Score 
    Applied Definition
    Numerical score indicating truthfulness of model in an interaction.
    AI RMF Characteristic(s)
    Safe
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Safety Evaluation
    References
    1ref
    • Su, Zhe, Xuhui Zhou, Sanketh Rangreji, et al. 2025. “AI-LIEDAR: Examine the trade-off between utility and truthfulness in LLM agents.” Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 11867–94. Link
  • Measurement Method
    Name
    TruthfulQA 
    Applied Definition
    Benchmark that tests whether models give correct answers in a question and answer setting. (Primarily an open benchmark.)
    AI RMF Characteristic(s)
    Valid & ReliableSafeAccountable & TransparentAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Lin, Stephanie, Jacob Hilton, and Owain Evans. "TruthfulQA: Measuring how models mimic human falsehoods." In Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers), pp. 3214-3252. 2022.
  • Measurement Method
    Name
    Usability and UX Research 
    Applied Definition
    Evaluation of efficiency, effectiveness, and user satisfaction in people’s interactions with AI systems. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Valid & ReliableSafeSecure & ResilientAccountable & TransparentExplainable & InterpretablePrivacy-EnhancedFairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Human-Centered Evaluation
    References
    4refs
    • Amershi, Saleema, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh et al. "Guidelines for human-AI interaction." In Proceedings of the 2019 CHI conference on human factors in computing systems, pp. 1-13. 2019
    • Microsoft HAX Toolkit
    • Google People + AI Guidebook
    • IBM Design For AI.
  • Measurement Method
    Name
    User Surveys 
    Applied Definition
    Structured questionnaires used to collect standardized feedback from users or stakeholders about an AI system. Informed consent, data protection, and legal or ethical approvals may be required for any human-subjects research involving users, stakeholders, or other participants.
    AI RMF Characteristic(s)
    Valid & ReliableSafeSecure & ResilientAccountable & TransparentExplainable & InterpretablePrivacy-EnhancedFairAccurate & Bias-Managed
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Human-Centered Evaluation
    References
    3refs
    • Ameller, David, Claudia Ayala, Jordi Cabot, and Xavier Franch. "How do software architects consider non-functional requirements: An exploratory study." In 2012 20th IEEE international requirements engineering conference (RE), pp. 41-50. IEEE, 2012
    • Borsci, Simone, Alessio Malizia, Martin Schmettow, Frank Van Der Velde, Gunay Tariverdiyeva, Divyaa Balaji, and Alan Chamberlain. "The chatbot usability scale: the design and pilot of a usability scale for interaction with AI-based conversational agents." Personal and ubiquitous computing 26, no. 1 (2022): 95-119
    • Amironesei, Razvan, Afzal Godil, Craig Greenberg, Kristen Greene, Theodore Jensen, Patrick Hall et. al. "Assessing Risks and Impacts of AI (ARIA) ARIA 0.1: Pilot Evaluation Report" NIST AI 700-2, Gaithersburg, MD, USA (2025).
  • Metric
    Name
    Verbatim Plagiarism Rate 
    Applied Definition
    Percentage of documents that LLMs generate verbatim, with exact overlap between LLM model generation and ground-truth text.
    AI RMF Characteristic(s)
    Secure & Resilient
    AI Lifecycle Stage(s)
    Collect & Process DataBuildUseOperate & Monitor
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    3refs
    • Lee, Jooyoung, Thai Le, Jinghui Chen, and Dongwon Lee. 2023. “Do Language Models Plagiarize?” Proceedings of the ACM Web Conference 2023, 3637–47. Link
    • Chen, Tong, Faeze Brahman, Jiacheng Liu, et al. 2025. ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-Training Data. Link
    • Lee, Jooyoung, Toshini Agrawal, Adaku Uchendu, Thai Le, Jinghui Chen, and Dongwon Lee. 2025. “PlagBench: Exploring the Duality of Large Language Models in Plagiarism Generation and Detection.” In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), edited by Luis Chiruzzo, Alan Ritter, and Lu Wang. Association for Computational Linguistics. Link
  • Measurement Method
    Name
    WMDP 
    Applied Definition
    Benchmark for hazardous knowledge and dual-use capability. (Primarily an open benchmark, with safety filters.)
    AI RMF Characteristic(s)
    Safe
    AI Lifecycle Stage(s)
    BuildUseOperate & Monitor
    Primary TEVV Application
    Language Model Benchmark Evaluation
    References
    1ref
    • Li, Nathaniel, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D. Li et al. "The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning." Proceedings of Machine Learning Research 235 (2024): 28525-28550.
  • Measurement Method
    Name
    Zero-Knowledge Proof 
    Applied Definition
    Cryptographic verification that a statement is true without providing any additional information; presence of verification enhances privacy.
    AI RMF Characteristic(s)
    Privacy-Enhanced
    AI Lifecycle Stage(s)
    Plan & DesignBuildUseDeploy
    Primary TEVV Application
    Privacy & Security Evaluation
    References
    2refs
    • Loic Lesavre, Priam Varin, and Dylan Yaga. Blockchain Networks: Token Design and Management Overview. Technical Report NIST Internal or Interagency Report (NISTIR) 8301, National Institute of Standards and Technology, February 2021. Link
    • Chuan Zhao, Shengnan Zhao, Minghao Zhao, Zhenxiang Chen, Chong-Zhi Gao, Hongwei Li, and Yu-an Tan. Secure Multi-Party Computation: Theory, practice and applications. Information Sciences, 476:357–372, February 2019. ISSN 0020-0255. doi: 10.1016/j.ins.2018.10.024. Link

Community feedback

The AI Metrology Center is a living resource and is expected to evolve over time. Individuals are encouraged to provide feedback about the content of the Center by emailing aimetrologycenter@nist.gov; we particularly welcome feedback on the filter categories, such as "Primary TEVV Application" or "TEVV Instrument." AI Metrology Center content updates will be released periodically.

Disclaimer

This collaborative resource is hosted by NIST. Content reflects contributions from multiple parties and does not imply endorsement or agreement by NIST. References to commercial products, systems, or trade names are for identification only and do not constitute recommendation or preference. Information is provided "as is," without warranties of any kind, including accuracy, completeness, merchantability, fitness for a particular purpose, or non-infringement.