Practigence System Integration Banner
Practice × Intelligence

In the AI era,
the power to ask questions and organize variables.

Moving from the "ownership" of knowledge and information to "processing"—how we use them.
Practigence is a next-generation assessment tool that visualizes the process of filtering cognitive noise and making essential decisions in today's rapidly changing world.

Practigence Cognitive and Natural Integration

Visualizing an Individual's True Processing Capacity (CPU) Based on Insights from Various Fields Beyond Cognitive Science, and the Developer's Expertise

Practigence objectively reveals "fluid intelligence" and "metacognitive ability," which traditional written assessments and multiple-choice tests fail to capture, aiming to optimize intellectual resources through placing the right talent in the right place.

🧠

From Static to Dynamic Intelligence

Rather than measuring Raymond Cattell's "crystallized intelligence (accumulation of knowledge)," it precisely measures "fluid intelligence"—the capacity to identify dynamics among variables from incomplete information and dynamically update hypotheses under rapidly changing situations.

🤖

"De-cognitive Saturation" in the AI Coexistence Era

As information supply becomes effortless due to generative AI, this tool evaluates "information filtering capacity" and "task formulation capacity" to discard cognitive noise and select only essential variables, keeping in mind the limitations of brain working memory (Miller's Law).

📊

Objective Visualization of Thinking Processes

It tracks not only final answers but also response latency (thinking time) and modification logs (Process Tracing). Integrating multi-faceted evaluation engines achieves fair assessment unaffected by the cognitive ceilings of evaluators (Ceiling Effect).

[Interactive Simulator] CPU (Fluid Intelligence) vs HDD (Knowledge Accumulation)

Simulate how differences in brain "processing structures" yield different outcomes in real-world problem-solving.

HDD (Knowledge Accumulator) Processing

// Waiting for boot...

CPU (Fluid Intelligence Executor) Processing

// Waiting for boot...

Philosophy

Promoting the Right Talent in the Right Place in Society

Practigence aims to promote the true placement of the right talent in the right place, optimizing intellectual resources in society. By establishing an environment where individuals can accurately grasp their own objective cognitive structures (CPU characteristics) and make optimal, autonomous decisions, we eliminate inefficiencies in intellectual productivity and enhance society's overall problem-solving capability.

Definition of "Intelligence" in Practigence

A dynamic, hierarchical cognitive processing capacity that metacognitively evaluates environmental suitability, autonomously selects, constructs, or discards environments to adapt to, extracts and reconstructs invariant logical structures from experience to understand core essences, models invisible causal dynamics as systems, flexibly updates hypotheses (Bayesian updating) even from minimal information, and simultaneously solves future risk prediction/prevention and current complex challenges.

Scientific Value Offered by Objective Assessment

As metacognitive theory in psychology (Flavell, 1979) shows, correct self-awareness of cognitive biases (self-calibration) requires advanced cognitive capability, and correcting biases solely by oneself can be difficult (Dunning-Kruger effect). Furthermore, because self-reported questionnaires trigger the "social desirability bias," leading respondents to unconsciously polish their answers, objective assessment free from subjective intervention becomes essential.

Objective measurement unravels thinking processes that tend to be treated as black boxes:

  • Visualization of Cognitive Balance: Logically organizes what combination of modules yields specific recognition and judgment characteristics.
  • Cheat-Proof Assessment Design: Captures true adaptability not through superficial correctness or self-selection, but by analyzing constraint classification, variable sorting, and response update logs along the timeline.
  • Improving Decision Accuracy and Autonomy: Grasping the characteristics and tendencies of one's own fluid intelligence enables "collaborative intelligence design" to appropriately leverage others and AI, supporting highly accurate decision-making.

The Fallacy That "Mathematics Encompasses Logic"

There is a common discourse in defense of education stating that "mathematics cultivates logical thinking." However, problems in mathematics are mostly closed environments where conditions are explicitly stated and irrelevant information is removed. In contrast, real-world decision-making deals with incomplete variables and varying information freshness.

While saying "because someone is good at math, they must possess logical thinking" is inappropriate, saying "because someone has logical thinking, they will likely be good at math" is appropriate. Mathematics is not logic itself; rather, logic encompasses mathematics.

Verbal Logic and Social Interference

What humans use to make decisions and communicate in society is vast amounts of "verbal logic" and only a tiny fraction of "mathematical logic." Despite this, overemphasizing mathematical logic just because it is easier to measure in school tests is inappropriate.

Against the correlation fallacy that equates good academic grades with high intelligence, Practigence provides a perspective to look straight at the contents and properly analyze and evaluate psychological biases where correlation fails.

Introduction

This tool provided by Practigence is designed to evaluate fluid intelligence based on a framework that decomposes and classifies cognitive processing required for real-world problem-solving into functional units.

This tool organizes what cognitive, judgment, or resolution failure occurs when specific functions are lacking, translates each capability into a measurable format, and conducts measurements using that framework.

Based on this definition, we perform functional decomposition to break down capabilities, devise and create methods to draw out, verify, and analyze whether those sub-items are successfully demonstrated, and generate reports from the collected data.

[Example] Functional Decomposition and Cognitive Load in Information Selection

To explain the decision-making process in daily life and business, let's consider a familiar scenario: **"purchasing a sofa."**

When considering a sofa purchase, one typically reviews basic variables such as "dimensions to fit the room," "price," "durability," and "color." However, if the attribute of a "single person who relocates frequently" is added, "ease of disassembly and transport" and "resale value" turn into "essential variables" for the decision. Conversely, for someone with an extremely spacious room, constraint variables regarding dimensions become irrelevant "noise."

As Cognitive Load Theory shows, our brain's working memory capacity is limited. If unnecessary information (noise) cannot be discarded from the head, processing resources required for judgment saturate, causing decision freezes or judgment errors.

Practigence evaluates this dynamic **information selection & functional decomposition ability (CD)** to determine "which variables are essential and which are discarded as noise in response to shifting premises."

[Interactive Simulator] Method of Functional Decomposition (1): Sofa Purchase Thought Experiment

In the purchase scenario explained above, click on different buyer attributes to visualize which variables become essential and which become unnecessary "noise."

Click on a buyer's "attribute" to see changes in variables to consider and unnecessary variables (noise):

Standard purchasing process: Decides by considering room and sofa dimensions, price, and delivery timing. At this stage, standard variable processing capability is required.

Required
Sofa Dimensions

Basic variable to check if it fits the placement area.

Required
Price & Durability

Practical variables relative to budget and usage period.

Required
Ease of Disassembly/Transport

Important constraint condition that arises when relocation is assumed.

Required
Resale Value

Variable to prevent back-end loss when letting go after short-term use.

Noise (Unnecessary)
Other People's Room Layouts

Layout information of other rooms with completely different shapes is noise that misleads judgment.

Required
Room Layout/Dimensions

Essential in standard placement planning, but its consideration weight drops to zero in extremely large rooms.

[Explanation] Method of Functional Decomposition (2): Decision-Making During Sudden Train Suspension

<Assumed Situation>
At an unfamiliar place, Train A (20-minute ride) heading to your destination suddenly suspends service due to track intrusion. The departure of the bus you must catch at the destination (absolute deadline) is in 30 minutes. You have only 2 minutes to decide.

Under this tense situation, a person with high dynamic intelligence executes the following "variable classification" and "high-speed computation" in their head instantly.

Option Presented Surface Info Variable Classification & Processing in Brain Computation Result
Route A: Train (20 mins) Currently suspended. May resume shortly. Deems "track intrusion" as a variable parameter with maximum volatility, ranging from minutes to hours depending on police intervention. Excludes immediately due to high risk against the 30-minute absolute deadline. Immediately Rejected (Cut Loss)
Route B: Bus + Train Bus arrives in 20 seconds, but route is uncertain and slow. The 20-second window physically fails because "recognition and movement costs" in an unfamiliar place exceed "available time." Also, travel time is long (noise cut). Immediately Rejected (Noise Filtering)
Route C: Subway (25 mins) Travel time is 25 minutes against 30 minutes remaining. Determines that the subway, structurally less affected by external ground noise, has 25 minutes of travel time as the "most reliable fixed parameter." Calculates that a 5-minute buffer is adequate to absorb transfer loss. Confirmed (in 2 mins)

Recruiter Cognitive Limits and "Process Black-Boxing":
The essence of this decision-making is cool-headed expectation value optimization: choosing the option with the highest probability of fitting within the deadline (30 mins), even if it means discarding the fastest (20 mins) possibility.

Non-linearity of Experience and the Gap in "Extraction Capacity"

Generally, the word "experience" is misunderstood as something whose value increases in proportion to time spent. However, length of experience (accumulation along the timeline) merely proves the physical fact of staying in that environment.

Time spent in uncritical inertia, without changing anything in the system or one's own actions, does not contribute to updating intelligence, no matter how vast it is.

The quality and speed of learning—how much one learns from experience—depends not on time length, but on how much pure logical structure or essential variables one can [extract] from that experience (concrete events).

Static Accumulation Experience

Accumulates concrete, surface-level events as-is (HDD model). Since context-dependent knowledge is stored as static memory, it cannot be applied when variables or environments change.

  • Characteristics: Dependent on past successful patterns, weak to changes in premises.
  • Evaluation Limitation: Interviews and traditional tests measure this accumulation, failing to detect vulnerability in chaotic environments.

Dynamic Processing Experience

Extracts abstract, invariant logical structures from concrete events (CPU model). By generalizing experiences as functions, it allows immediate application and hypothesis updating in unknown, novel tasks.

  • Characteristics: High structural transfer (analogical transfer) capacity, flexible hypothesis updating.
  • Observation Goal: Practigence evaluates this extraction and transfer process through variable constraints and AI dialogue logs.

Assessment Output: Report Tiers

Data computed after taking the assessment is provided as reports in three grades depending on use and objectives (currently, Tier 1 and Tier 2 are accessible to examinees).

  • Tier 1 — Behavioral Description Summary (Free Plan / Open to Examinee)

    A simplified report describing the specific movements of cognitive processes observed during the test in text. It provides 5 to 7 lines of behavioral observation logs without scores or tags, such as: "tended to rush to conclusions in multiple tasks" or "integrated new information while retaining the initial hypothesis" (analogous to recording "shallow sleep this week" in a health checkup context).

  • Tier 2 — Cognitive Cluster Report (Paid Plan / Open to Examinee)

    Our core report showing levels, analytical comments, strengths, and weaknesses across "13 cognitive clusters" (such as logical analysis, reasoning, and verbal transmission), along with specific recommended improvement actions like "asking specific questions immediately after daily decision-making" (analogous to functional evaluation and guidance like "Cardiopulmonary: Good / Sleep: Needs Improvement" in a health checkup context).

  • Tier 3 — Precision Detailed Report (Premium Grade / Open to Admins Only)

    Detailed data for professional analysis mapping presence/absence for all 80+ high-order cognitive atomic tags per task, including pattern analysis of how cognitive structures collapse under specific question types, and automatically computed tags requiring prioritized improvement (analogous to a complete list of blood test values used for research and detailed analysis).

Features & Ability Categories

Practigence comprehensively decomposes thinking processes from the moment a human "recognizes" things to "judgment/decision-making," defining and measuring them in domains A to N, transcendent ability M, and holistic integration P.

Behavior-Centric Evaluation: Misidentifying "High Intelligence" as "Developmental Disorder"

Because traditional checklist evaluations measure superficial behaviors (What) rather than individual processing mechanisms, they run the risk of misidentifying highly intelligent individuals as having "developmental disorders." Click tabs to compare.

Superficial Behavior (What: judged as ADHD/Inattention)

Fails to sustain attention on specific tasks due to difficulties in working memory control, unintentionally looking at other stimuli. Appears unable to maintain attention when facing boring tasks.

Processing in Brain (Why: surplus resources from high intelligence)

Because the brain's processing speed is too fast, simple tasks are completed instantly, leaving working memory mostly vacant. Consequently, to maintain the brain's optimal arousal level, the individual voluntarily starts separate, highly demanding thoughts or intellectual exploration (mental multitasking). This is not "inattention," but self-protective optimization to prevent brain stall. Shows extreme focus on difficult tasks.

Superficial Behavior (What: judged as ASD/Social Disability)

Difficulty deciphering other people's intentions or unwritten social contexts (like reading the room). Does not follow irrational rules. Thought processes skip steps and appear odd. Ignores implicit rules.

Processing in Brain (Why: mismatch in cognitive levels)

Because premises, levels of abstraction, and logic processing speed differ too much from counterparts, the individual's sound reasoning or essential arguments are not understood by peers. Non-compliance with rules is because the "uselessness of compliance" is too logically clear. Skip-step reasoning makes conclusions appear "abrupt and odd" to those around them.

[Cognitive Science Perspective] Evaluator Cognitive Ceiling and Misclassification of Traits

As the Law of Requisite Variety (Ashby, 1956) in cybernetics shows, evaluating and controlling a complex system requires the evaluating side to possess equal or greater variety (complexity of states). If there is a massive gap between the evaluator's (manager, teacher) cognitive traits and the examinee's, the evaluator may fail to recognize the examinee's advanced processing or meta-perspectives, causing hiring or management mismatches.

Thus, actions that are actually "voluntary exploration to avoid boredom" or "custom variable design" are superficially mislabeled as "inattention" or "lack of cooperation" due to the structural limitations of conventional systems.

Need for Multidimensional Intelligence Models:
The dominance of knowledge (crystallized intelligence) in teaching and the dominance of processing (fluid intelligence) in solving unknown tasks lie on different dimensions. Rather than evaluating superficial behavior (What) like traditional tests, modeling behavior as an interaction (function) of "internal resources (CPU)" and "external task load (premises)" excludes misclassifications from cognitive limits, capturing the individual's essential cognitive traits objectively.

Transition to an Evaluation System Applying Dynamic Assessment

Conventional System (Limit of Single-Output Structure)

Directly classifies the single output of surfaced "behavior," ignoring underlying contexts and variables.

  • Input: Behavior (e.g., doing something else without focusing on class or work)
  • Processing: Deems the behavior as "inattention (noise)" and immediately labels it as "suspected ADHD (lazy)."
  • Issue: Decisive dynamic variables such as an individual's "internal resources (CPU)" and "environmental/task load" are completely ignored (or treated as constants).

Assessment System of This Tool (Dynamic Variable Classification)

Rather than directly classifying superficial behavior, it properly factorizes it as an interaction (function) between "internal resources" and "external environment."

  • Identify Internal Resources (Fixed Parameters): Accurately grasp an individual's pure "brain processing speed" and "working memory capacity" via assessment as unchangeable premises.
  • Match with Task Load (Variable Parameters): Objectively evaluate the difficulty of the presented environment and the noise level of information.
  • Rational Judgment: If environmental load is too low relative to the individual's resources, starting separate thoughts is processed as an "inevitable and rational consequence," eliminating misclassification labels.

Ability Category System (A - P)

A: Recognition Domain

What was recognized and perceived

When facing a situation, which variables or noises are first caught intuitively. This is the starting point of all cognitive processing. Failures in recognition at this level directly lead to subsequent analysis or judgment errors.

Features & Value

🔬

Multi-Faceted AI Profiling

Fine-Grained thought Process Visualization

Multiple AI evaluation engines analyze your answers independently. Instead of simple pass/fail or correctness, they map and visualize the micro-level thought processes, including where logic leaps or hypotheses are revised.

🧠

Knowledge-Independent Design

Pure Measurement of Thinking Capacity

Unlike typical tests that require memorized facts, it evaluates your capacity to identify variables and infer logic when facing unfamiliar problems. This ensures no bias based on academic background or major.

⚖️

Fair Weighting by Difficulty

Sophisticated Scoring Weights

Rather than scoring easy mistakes and hard successes equally, scoring is weighted according to task complexity and variable density. This accurately evaluates the examinee's true processing boundary.

🎯

Structural Design for Accuracy

Countering Bias and Skewed Judgments

To capture outstanding talent without oversight, prompt parameters and sequences are structured to avoid evaluation bias. This guarantees an objective, fair, and highly accurate assessment.

🛡️

Resistance to Test-Prep & Gaming

Dynamic Item Variations

Key variables change dynamically with each session, making memorized answers or test-prep techniques completely useless. The test measures real-time processing rather than static memory.

🔒

Privacy & Autonomy Focus

No Personal Data Mapping & Cooldowns

Test results are not linked to identifiable personal information to secure privacy. Additionally, a strict cooldown interval is enforced before retakes to eliminate simple practice effects.

"Advantage" for Whom?

Advantage for Individuals

Self-understanding of cognitive traits and guidelines for cultivating future intelligence

Enables objective calibration of one's own thinking habits, such as: "in what situations are judgment biases likely to occur" and "which information failed to be processed as noise." Going beyond merely revealing current strengths and weaknesses, it provides concrete guidelines on how to foster irreplaceable high-order cognitive skills unique to humans in the AI coexistence era.

Guidelines for intellectual growth to break away from knowledge-reproduction models
Self-calibration to discover biases in thinking processes

Advantage for Enterprises

Measuring processing capacity (CPU) adaptable to the AI era and preventing hiring mismatches

Evaluates pure thinking and adaptability (CPU) to tackle unknown tasks, rather than volume of memorized knowledge (HDD). Because the structure makes it difficult to fake or prepare in advance, recruiters can objectively grasp "true cognitive processing capacity"—which is hard to evaluate through interviews alone—drastically reducing organizational costs associated with hiring and placement mismatches.

Direct measurement of CPU-type processing capacity to deal with chaotic environments
Process-tracking assessment that invalidates preparation and exaggeration

Advantage for Society

Upgrading evaluations from knowledge-overemphasis to cognitive processing, and promoting the right talent in the right place

Upgrading the conventional evaluation infrastructure—which only favored those who remembered a lot of past knowledge (crystallized intelligence)—to center on the "high-order fluid intelligence" required in the highly volatile and unpredictable AI era. We promote the right placement of cognitive resources according to each individual's diverse cognitive structure, encouraging a collaborative society where individuals leverage strengths to complement one another autonomously.

Essential intelligence evaluation infrastructure that remains sustainable in the AI coexistence era
Placing the right cognitive resources in the right place, leveraging cognitive diversity

Test Guide: How to Take the Test

Target: Prospective Examinees

This test is not a "trivia quiz to guess the correct answer." We look at **how you processed information and made judgments**. Therefore, rather than trying to write a polished answer, it is most important to **express your thought process in words as it is**. Reading the following before starting will help you demonstrate your capabilities to the fullest.

Steps of Assessment

There is a Consent Gate prior to testing. The test has a 2-part structure (Part 1 and Part 2) and can be safely paused and resumed at any time.

48-Hour Limit

Part 2 must be taken within a certain period (48 hours) after completing Part 1. The expiration time is clearly indicated on the screen.

Process-Extracting Items

Rather than fixed memorization items, tasks are designed to extract your thought process. Some items change dynamically per session, rendering test-prep ineffective.

Detailed Profiles

Provides your level from L0 (base unreached) to L11 (flawless), profiles by field, and a breakdown of "what is working / what is not."

Assessment Workflow

  1. Consent Gate: Complete the evaluation item consent procedure before the test.
  2. Part 1 Tasks: Even if you pause, your answers up to that point are safely autosaved and can be resumed later.
  3. Completion of Part 1 & Interval: Take Part 2 within 48 hours of finishing Part 1. The expiration time is displayed on the screen.
  4. Part 2 Tasks: Answer the remaining questions. Dynamic items are included, testing your raw, real-time reasoning.
  5. Results: Results are detailed as L0-L11 levels, profiles by field, and "what is working / what is not."

Rules & Guidelines

  • Revision RightUp to **3 modifications** to your answers are allowed. However, some questions are **not editable**, which will be explicitly stated before the question.
  • Copy/PasteIn principle, pasting is prohibited. As an exception, **copying & pasting is permitted only for information processing tasks**, which is explicitly indicated above the response field.
  • All Fields RequiredFor questions with multiple sub-questions, **you cannot proceed to the next page unless you answer all of them** (a warning will be displayed).
  • Behavior TrackingSwitching tabs or windows, attempting to paste, taking screenshots, or print operations will be **recorded**.
  • Long Passage UXLong passages are displayed **from the very top**, and response fields can be **collapsed**. This allows you to focus on the text before writing.

4 Tips to Demonstrate Your Capability Properly

  • Your Written Answer is Everything. Scoring is done solely based on what you write. Thoughts remaining in your head cannot be evaluated unless written down.
  • Write Hypotheses Even if Unsure. Do not stop at "I don't know." We evaluate your attitude in putting forward a tentative answer backed by reasoning.
  • Verbalize Your Reasoning & Process. Always include "why you judged so" in addition to your conclusions.
  • Preparation and Memorization are Futile. The test is designed to measure on-the-spot processing, not memorized answers. Approaching it raw is best.

Levels Concept & Interpreting Results

Results are indicated as levels from L0 (base unreached) to L11 (flawless). This level is not based on a single sharp skill, but is integrated such that you only advance when Perception, Analysis, Thinking, Design, Objectivity, and Verbalization all function simultaneously. At the highest tiers, processing speed is also included in the conditions.

The report clearly distinguishes between "active skills" (currently functioning stable strength) and "latent potential/talent" (inspiration or possibility that is not yet stable) using graded labels. Please use this not as a label of superiority but as a guide for your cognitive evolution and future learning.

Methodology & Validity

Target: Experts, Peer Reviewers, Technical Evaluators

This page describes the construct concepts of the measurement, measurement models, scoring architecture, reliability assurance, and the **arguments and limitations of validity at this point** without exaggeration. We distinguish between what is achieved and what remains unestablished in a verifiable manner.

1. Constructs of Measurement

The primary measurement target is **fluid intelligence**—the ability to identify relationships and structures from novel, incomplete information and dynamically update hypotheses independent of acquired knowledge. Along with this, we evaluate **self-objectivity (calibration) based on metacognition**. The theoretical background of these constructs, the mapping of each field to research constructs, and references are aggregated in the "Theoretical & Research Basis" page.

2. Measurement Model Concept: Functional Decomposition to Hierarchical Levels

Cognitive processing required for real-world decision-making is **decomposed into functional units**, systemized so that "which cognitive, judgment, or resolution failure occurs when specific functions are lacking" can be backward-computed. Decomposed functions are aggregated into cognitive clusters, and attainment is expressed in progressive levels. The targets of observation are assigned not based on "given question formats" but starting from **what can be observed in each task (observability)**; the mapping between target constructs and tasks is traceable by design.

3. Scoring Approach: Cross-task Evaluation, Multi-model Consensus, Higher-order Arbitration

Scoring is not done per question; instead, evaluation criteria are judged **across all responses**, suppressing localized judgment volatility. Judgments are made by **consensus among multiple independent large language models**, and only instances where models disagree are arbitrated by a higher-order model. This architecture aligns with findings in evaluation research showing that ensemble (consensus) and reasoning-explicit designs yield higher consistency and transparency than single models.

Cross-task
Evaluation per Criterion

Each evaluation criterion is assessed across all responses, applying reasoning and confidence levels.

Overview
Holistic Evaluation

Evaluates qualities and flaws as a whole that cannot be reduced to individual criteria.

Consensus
Multi-model Parallel Run

Multiple independent models score and cross-reference results.

Arbitration
Higher Arbitration on Disagreement

Higher-order model arbitrates only where judgments diverge.

* Designed to route tasks by degree of disagreement: areas resolved by consensus are processed lightly, while only split-decision points are forwarded to higher-order arbitration (balancing cost and accuracy).

4. Ensuring Reliability

  • **Multi-model Consensus:** Smooths out accidental fluctuations of a single evaluator by cross-referencing multiple independent models.
  • **Explicit Reasoning:** Enhances consistency and transparency of evaluations by requiring explanations for each judgment.
  • **Denominator Correction:** If observation opportunities for a criterion are insufficient, evaluates the validity of the omission to correct reliability.
  • **Process Tracing:** Records behavioral metrics such as response latency (thinking time), edit logs, tab switching, and paste attempts as auxiliary info.
  • **Resistance to Gaming/Manipulation:** The scoring logic is stored solely server-side, and responses are treated as "data" with injection prevention to block overwrite/instruction injection attempts by examinees.

5. Arguments for Validity (Content-related & Construct-related)

**Content Validity:** Each task has evaluation criteria assigned based on observability; design-wise, the mapping between target constructs and tasks is traceable. **Construct Validity:** The hierarchy of functional decomposition -> clusters -> levels is structured to correspond to the theoretical frameworks of fluid intelligence and metacognition. It intends to capture the "quality of processing" that existing metrics (which measure only outcomes) miss, by observing processes. For details on correspondence with theory and research, see the "Theoretical & Research Basis" page.

6. Current Limitations and Future Verifications (Honest Disclosure)

The system's design was established first, and verification/research comparison was conducted post-hoc to support and confirm its correctness. This evaluation is based on observing the thought process, measuring fluid intelligence (dynamic, CPU-like intellect), and does not define a person's human value. With these premises in mind, we honestly disclose the following limitations:

  • **Probabilistic Variation of LLM Scoring:** While mitigated by multi-model consensus and higher-order arbitration, fluctuations cannot be entirely eliminated since the evaluators are probabilistic. Especially in judgments requiring specialized knowledge, it is known that the agreement rate between humans and AI drops compared to general tasks. Note that while we minimize probabilistic variations through prompt engineering, system design, role definition, and buffers, there remains a possibility of approximately 0.125% that partially incorrect content may be generated.
  • **Lack of Standardization/Norms:** Standardizing (norming) across large populations and empirically calibrating level boundaries remain future challenges. The current levels are theory-driven criteria.
  • **Unverified External Validity:** Correlations with work performance or other established metrics (criterion-related/predictive validity) are at a stage where they must be confirmed through replication studies using external criteria.
  • **Test-Retest Reliability:** Stability over time during retesting (test-retest) requires separate measurement.
  • **Scope of Application:** This tool aims to map cognitive traits and is not a clinical diagnostic tool. Descriptions regarding developmental traits represent "explanations of misidentification structures" and do not constitute a diagnosis.
  • **Linguistic & Cultural Scope:** Currently, the tool is primarily designed for Japanese speakers, and equivalence across languages and cultures remains unverified.
To Theoretical & Research Basis → Take the Test

Theoretical & Research Basis and References

This page aggregates the theoretical and research foundations with which Practigence's **intelligence framework (functional decomposition of abilities) and test design** align, demonstrating how each field maps to established research constructs and how alignment with the references was verified.

Positioning of References

**Design came first; verification, validation, and refutation matching were performed afterwards.** The research, findings, and refutation studies mentioned on this page were all **used to assist and verify the accuracy of the designer's own conceptualization, design, and hypotheses**. Rather than using external findings as a starting point or adopting them as conclusions, they were referenced post-hoc as comparison materials to confirm alignment between already constructed frameworks/test designs and hypotheses. Therefore, references on this page are not intended as replication, replacement, or borrowing authority from each study.

A. Theoretical Foundations of Constructs

The primary measurement targets are **fluid intelligence**—the ability to identify relationships and structures from novel, incomplete information and dynamically update hypotheses independent of acquired knowledge—and **metacognition (self-objectivity/calibration)** to monitor and control one's own cognition. These have been confirmed to align with the following lineages of research:

  • The concept of fluid intelligence as distinguished from crystallized intelligence (Cattell, 1963), and the subsequent hierarchical theory of cognitive abilities (Carroll, 1993 / CHC Theory).
  • The framework of metacognition that monitors and controls one's own cognition (Flavell, 1979 / Monitoring-Control by Nelson & Narens), and systematic biases in self-evaluation (Dunning & Kruger, 1999).
  • Limitation of working memory capacity (Miller, 1956 / Baddeley & Hitch, 1974) and Cognitive Load Theory (Sweller).
  • Deliberate processing in Dual-Process Theory (Kahneman, 2011).
  • The Law of Requisite Variety (Ashby, 1956), stating that an evaluating system requires variety equal to or greater than the target.

B. Correspondence Between Intelligence Framework and Research Constructs

Each field in this framework is not an arbitrary division, but corresponds to constructs repeatedly identified in cognitive science and intelligence research. The validity of the framework is supported by these correspondences.

Framework FieldCorresponding Research ConstructPrimary Basis
A Recognition / CD SelectionSelective Attention, Relevance Filtering, Cognitive Load ManagementCognitive Load Theory (Sweller) / Attention Research
B Analysis / C UnderstandingRelational Inference, Structure Mapping, Inductive Reasoning (Core Gf)Structure Mapping Theory (Gentner) / Cattell–Horn–Carroll
D Memory / Ds Selective MemoryWorking Memory, Abstraction & Analogical TransferBaddeley & Hitch (1974) / Analogical Transfer (Gick & Holyoak)
E ThoughtFluid Reasoning, Hypothesis Verification, Bayesian Belief UpdatingCattell (1963) / Bayesian Cognition (Tenenbaum et al.)
F Design/Solution / G Social ImplementationIll-Defined Problem Solving, Practical IntelligenceProblem Solving Research / Sternberg's Practical Intelligence
K Self-Objectivity / L Holistic Objectivity / I Env SelectionMetacognition, Self-Regulation, Executive FunctionFlavell (1979) / Nelson & Narens
J Cognitive EnduranceSustained Attention, Cognitive DurabilitySustained Attention Research / Recent Empirical Evidence on Cognitive Endurance
M Fusion / P HolisticGeneral Intelligence Factor g, Integration of Abilities, Successful IntelligenceSpearman's g / Carroll (1993) / Sternberg's Successful Intelligence

**The validity of centering on "fluid intelligence × metacognition"** is also reinforced by recent intelligence research. Fluid reasoning is stably identified as the core of the CHC hierarchy, and metacognition/executive functions are positioned as independent predictors of learning outcomes and adaptive judgments. * This indicates correspondences at the field level.

C. Empirical Support for Societal Necessity

Following the widespread adoption of generative AI, empirical studies show a decline in the relative value of knowledge-reproducing tasks and an **increased demand for high-order cognitive skills**. Jobs that assume the use of generative AI tools exhibit significantly higher cognitive skill requirements (Hampole et al., 2025, arXiv:2503.09212), and analyses by the IMF and McKinsey also highlight impacts on knowledge work and a **supply shortage of critical thinking, problem structuring, and complex information processing** (IMF Staff Discussion Note, 2024 / McKinsey Global Institute). The constructs measured by this tool respond directly to this supply shortage area.

D. Evaluation Research Supporting Scoring Design

The design philosophy of using "consensus among multiple models" and "requiring explicit reasoning" in scoring aligns with recent LLM evaluation research. It has been reported that ensemble (consensus/majority vote) and explanation-requiring designs yield higher agreement and transparency than single models (Zheng et al., 2023 / Research on Reliability of Multilingual Evaluation, 2025). At the same time, it has been demonstrated that in judgments requiring specialized expertise, the agreement rate between humans and AI drops compared to general tasks (falling to 64–68% in specialized domains compared to ~90% human-AI agreement in general tasks, which is lower than the inter-expert agreement of ~72–75%), serving as the basis for this tool's recognition of its own limitations.

E. References and Their Application in the Framework & Tests

ReferenceApplication/Correspondence in Framework & Test
Cattell (1963) / Carroll (1993, CHC)Basis for positioning fluid intelligence as the primary target of measurement. Conceptual background of level hierarchies.
Flavell (1979) / Nelson & NarensMeasurement basis for metacognition and self-objectivity (Domains K/L).
Dunning & Kruger (1999)Basis for the design measuring gaps between self-declaration (pre/post-test) and actual measurements.
Miller (1956) / Baddeley & Hitch (1974)Cognitive load design for memory (D) tasks, noise filtering, and information overload tasks.
Sweller (Cognitive Load Theory)Design of cognitive load in long-form information processing questions.
Ashby (1956, Requisite Variety)Theoretical explanation of the evaluator-ceiling problem, motivation for multi-model consensus.
Kahneman (2011, Dual Process)Perspective for capturing validation of intuition and deliberate processing.
Gentner (Structure Mapping) / Gick & HolyoakDesign basis for analysis/understanding (B/C) and analogical transfer (Ds / transfer problems).
Sternberg (Practical/Successful Intelligence)Construct background for social implementation (G) and ability integration (P).
Schulte-Mecklenbeck et al. (2011)Methodological basis for process tracing (response latency, edit logs).
Zheng et al. (2023, LLM-as-a-Judge)Alignment with consensus and reasoning-explicit scoring designs.
Multilingual LLM-as-a-Judge (2025, arXiv:2505.12201)Supporting evidence for lower agreement rates in specialized domains = recognition of limitations.
Hampole et al. (2025, arXiv:2503.09212)Supporting evidence for increased demand for high-order cognitive skills = societal necessity.
IMF SDN (2024) / McKinsey MGIImpact on knowledge work and supply shortage of critical thinking & problem structuring.

* The above references are intended to position constructs, methods, and limitations; they do not assert that this tool is a replication or replacement of each study.

← To Methodology & Validity Home

Terms of Use & Consent Items

Last Revised: June 2026 / These Terms of Use define the conditions for using the Practigence Assessment (hereinafter "the Service"). By using (taking) the Service, examinees are deemed to have agreed to these Terms.

Article 1General Provisions & Applicability

These Terms apply to all relationships between the provider of the Service (hereinafter "the Provider") and examinees. Examinees must review and agree to these Terms and Consent Items before starting the assessment. If you do not agree, you cannot take the assessment.

Article 2Retake Policy (Cooldown)

The validity of measurement is compromised when examinees familiarize themselves with question formats and tasks (measurement contamination due to learning/practice effects). To prevent this, a specific cooldown period is established for retakes.

  • In principle, retakes by the same examinee are prohibited for **6 months** starting from the completion date of the initial assessment. (This period may be changed based on operational policies.)
  • Applications or assessments taken during the cooldown period will be invalid, and results will not be provided.
  • Retakes may be permitted at the administrator's discretion only in cases of system failure, provider-side reasons, or other legitimate grounds.
  • The act of taking the test using alternative accounts to bypass the cooldown constitutes a fraudulent activity under Article 3.

Article 3Prohibition of Fraudulent Activities

This Service measures "raw cognitive processing"; the following actions are prohibited. If violations are confirmed, measures such as invalidation of the assessment, non-provision of results, or suspension of use may be taken.

  • Creating or outsourcing answers using third parties, generative AI, search engines, etc. (except where explicitly permitted by the Service).
  • Proxy test-taking, using multiple accounts, bypassing cooldowns, or other acts that distort measurements.
  • Copying & pasting on unauthorized questions, referring to external materials, or other actions violating designated rules.

* Tab switching, paste attempts, and screenshot operations during the test are recorded and may be used for fraud detection. Collection of question data and reverse engineering are strictly defined in Article 3-2.

Article 3-2Prohibition of Question Data Collection & Complete Prohibition of Reverse Engineering

Questions, sub-questions, options, model answers, scoring logic, evaluation frameworks, tags/judgment criteria, judgment methods, model configurations, user interfaces, and all other components of the Service are intellectual property and trade secrets belonging to the Provider and legitimate rights holders. Examinees must not perform any of the following acts:

  • Recording, photographing, videotaping, taking screenshots, transcribing, copying, saving, storing, or creating datasets of question texts, sub-questions, options, model answers, scoring comments, report contents, etc. (hereinafter "Question Data").
  • Sending, disclosing, sharing, distributing, publishing, or redistributing Question Data to third parties, or supplying it for machine learning/AI training/evaluation or other uses.
  • Analyzing, decompiling, disassembling, guessing, restoring, reconstructing, or automatically acquiring (via scraping, bots, etc.) scoring logic, evaluation frameworks, tag definitions, judgment criteria, judgment methods, model configurations, thresholds, or other non-public elements.
  • Acts bypassing or evading the above, or causing or facilitating third parties to perform these acts.

In response to violations, the Provider may take legal action including injunction claims, damages claims, and criminal charges, in addition to invalidating the assessment and suspending Service use.

The validity and enforceability of specific provisions vary based on applicable law, jurisdiction, and individual circumstances. These constitute contractual agreements between the Provider and examinees; their finalization and operation require verification by an attorney. These Terms do not constitute legal advice.

Article 4Handling of Personal Information & Data

Basic Policy: The Provider does not collect or store assessment results associated with personally identifiable information. Result data is handled in a form that does not identify individuals (pseudonymization/identifier separation), and is not stored, matched, or provided to third parties in a way that links to a specific individual's profile.
  • Written and behavioral logs (thinking time, edit history, etc.) acquired during the assessment are used solely for scoring and report generation purposes.
  • Information temporarily required for access authentication, etc., is handled separately in a form that does not permanently link to assessment results.
  • The Provider will not attempt to re-identify individuals from results, nor transfer results to third parties without the individual's consent.
  • Even when data is used for statistical quality improvement and research, it assumes aggregation and anonymization that cannot identify individuals.

* This Article defines the Provider's handling policy. Details regarding specific retention periods and management frameworks follow a separately defined Privacy Notice.

Article 5Fees & Payment Matters

  • The Service contains areas provided free of charge and areas provided under paid plans (higher-grade reports, etc.).
  • Fees, billing units, payment methods, and scope of provision for paid plans follow conditions displayed at the time of application. Displayed and agreed conditions take precedence.
  • Due to the nature of digital services (scoring and report generation), refunds are in principle not provided after service delivery begins. Refund eligibility follows application conditions and relevant laws.
  • Payment of the assessment fee guarantees eligibility to take the test; it does not guarantee a specific score, level, or evaluation outcome.
  • Fees and plan details are subject to revision without notice. Revisions apply only to future applications.

* Specific amounts and plan details will be reflected in this Article once finalized.

Article 6Status of Results & Disclaimers

  • Results of the Service are intended for mapping cognitive traits and **do not constitute medical or clinical diagnoses**. Descriptions regarding developmental characteristics are generalizations for explanatory purposes; do not use them as bases for diagnosis or treatment.
  • Results are estimates under conditions at the time of scoring; since probabilistic evaluators (AI) are used, they contain certain fluctuations. The Provider does not guarantee complete accuracy of results or suitability for specific purposes.
  • The Provider bears no responsibility for the outcomes of decisions made by examinees or third parties regarding how they use the results (hiring, evaluation, self-judgment, etc.).

Examinees shall confirm and agree to each of the following matters before starting the assessment (consent is obtained for each item on the pre-test screen):

  • Do not engage in fraudulent activities (Article 3) and take the test in accordance with designated rules.
  • Understand and comply with the cooldown period for retakes (Article 2).
  • Understand the policy for handling personal information/data (Article 4) and consent to the collection/use of behavioral logs, etc.
  • Agree to the fees and payment conditions (Article 5) when using paid plans.
  • Understand the status of results and disclaimers (Article 6).
  • Agree to these Terms in their entirety.

Article 8Revision of Terms

The Provider may revise these Terms as necessary. Important changes will be announced on this page. Taking the test after revision constitutes agreement to the revised Terms.

← Home Proceed to Test