The Universal Baseline provides a minimum Evaluation Foundation that can be applied across AI Systems. However, in actual AI use, the same term, Output, Action, or Signal does not necessarily carry the same meaning or Risk in every situation. Its Significance may vary according to Context, including Purpose, User, Authority, Operation, Region, Language, and Culture.
When the Universal Baseline detects a potential Risk Signal, not every Case requires the same level of detailed Evaluation. Where necessary, the Case is connected to Contextual Evaluation through Internal Vigilance. Internal Vigilance examines Meaning, Intent, Use Context, Linguistic & Cultural Context, and other relevant factors to evaluate the Significance of the Signal in relation to the applicable Policy Objective.
Importantly, Contextual Evaluation ≠ Modification of the Policy Objective. Context does not alter the Policy Objective itself. Rather, it refines how the Requirement should be evaluated in a specific Situation.
For example, detecting the term “bomb” does not by itself determine the meaning of the Risk. Where a specific and significant Risk Signal can already be sufficiently established, additional Contextual Interpretation may not be necessary. Where Significance depends on Meaning or Intent, such as Quotation, Education, Historical Explanation, or Metaphor, additional Evaluation through Internal Vigilance may be required. Baseline Evaluation and Contextual Evaluation are therefore complementary rather than competing functions. The former provides a common Evaluation Foundation, while the latter enables deeper assessment of Meaning and Context where necessary.
A term such as “cocktail” further illustrates this Context dependency. In ordinary use, it may refer to a beverage and indicate no particular Risk, while as part of “Molotov cocktail,” it may contribute to a different Risk Signal. Across Languages, the same concept may be expressed through entirely different lexical and Semantic Structures. Words or surface-level Semantic Classification alone therefore cannot provide uniform Risk Evaluation across different Languages, Cultures, and Use Contexts.
The role of Internal Vigilance is to incorporate these Contextual differences into Evaluation and proceed from the Universal Baseline to Deeper Contextual Evaluation where necessary. However, the Depth, Method, Expertise Composition, and Technical Implementation of that Evaluation do not need to be fixed into a single approach. They may differ according to Provider, Organization, Domain, and Operational Context.
What should be standardized is not the Context itself, but the Vigilance Function that maintains the Universal Baseline, refines Evaluation through Context where necessary, and preserves Alignment with the higher-level Policy Objective.