Skip to content
Giordano Cabral

fluency · 2025–

What is fluency in artificial intelligence, and why measure movement rather than position?

Fluency is measuring movement instead of position. In a field whose state of the art changes every few months, a maturity scale places the organization on a reference that is already outdated. Movement is also not measured by what a person reports about themselves: the METR controlled study, from July 10, 2025, measured 16 experienced developers taking 19% longer with AI on 246 tasks from their own repositories, while they estimated they had gotten 20% faster. The measure is taken from the trace the work leaves. The concept underlies FAROL's six axes, PsyFun's games, and process-based assessment in Giordano Cabral's courses.

Workbench viewed from above: an MDF box with four buttons and a foot pedal, a cable, a spool of solder, a breadboard, calipers, and a sheet of paper with a wave drawn on it; a red button.

01 · Fluency

The METR study, 2025 and 2026

On July 10, 2025, METR published a controlled trial on productivity and AI. Sixteen experienced developers worked on 246 tasks in their own repositories, each randomly assigned to allow or prohibit AI use. With AI, they took 19% longer. At the end, they estimated that they had become 20% faster.

On February 24, 2026, the same authors repeated the study. Among the ten returning participants from the original group, the estimate was 18% faster, with an interval ranging from 38% faster to 9% slower — and the authors themselves call the new finding a weak signal, because a growing proportion of invitees decline a study that requires them to work without AI half the time.

The result flipped sign between the two measurements, and the experiment's design stopped measuring what it used to measure, because AI use among participants changed. A scale that assigns a rung assumes the scale itself remains stable.

The term adopted by Giordano Cabral is fluency. He has used it since 2025 at FAROL, a group he coordinates at CIn/UFPE with Filipe Calegário; before that, without this name, the same approach appears in PsyFun's games and in classroom assessment.

02 · Fluency

Why fluency

The term comes from language learning: without practice, the level regresses. Fluency is measured by axis — an organization can be fluent in data and not in people — and depends on the surrounding ecosystem.

Adoption alone no longer differentiates organizations: the 2025 Stack Overflow survey records 84% of respondents using or planning to use AI in development, against 76% the year before, and 46% who distrust the accuracy of what it returns, against 33% who trust it.

AI maturity models inherit from 1990s software maturity models, built for stable processes. FAROL's critique, detailed in agentic AI, has three points: the organization ranked at the top tends to have more to lose in the next shift; these models reward standardization and penalize experimentation, while in AI fast discard has value; and position does not indicate speed.

Six dimensions — knowledge, tools and integration, delegation and autonomy, processes, productivity and results, learning and culture — each with its own scale from N0 to N5, assessed against a master bank of 264 items. The entry point is the open diagnostic assessment: a leader answers on behalf of the organization in three to eight minutes, without prior registration, and receives the radar chart immediately. What a company does next is covered in business transformation with AI.

Five readings of the same bank of 264 items in the Mapa da Fluência Agêntica
Five readings of the same bank of 264 items in the Mapa da Fluência Agêntica

03 · Fluency

Measuring through traces

Asked whether they would cooperate, children almost always say yes. In PsyFun's games, cooperating costs a resource, the choice is made in about three seconds, and nothing on screen is called a dilemma. In two experiments with 162 children aged 6 to 12, published in Behavior Research Methods in 2021, they cooperated more in one game than in the other, and conditionally — two things a questionnaire does not separate. Cabral coordinated the group at UFRPE in 2013 and 2014 and, until 2018, the CNPq project that originated it.

The distinction is between what a person reports about themselves and the record of what they actually did. The two measures have errors of a different nature and, for that reason, do not add up: self-report errs systematically and in the same direction, as in METR; the trace only sees the work that passes through a system that logs it, and records what was done, not what was understood.

In Tendências em Mídia e Interação, 2026.2 edition, each student sifts through five hundred tools item by item inside the course system, which logs the time spent on each one. The rule is in the syllabus: if the machine produces the output, what gets graded is the process of piloting the machine. The course does not use an AI detector; the criterion is evidence of process — repository, history, justification, testing. In the first submission, fourteen students brought 5,776 tools to the shared catalog and logged 659 choices.

At FAROL, the picture of observed use is a separate dimension, outside the data collection ladder, because it is not added to what people report. Meu Mapa de IA asks students what they have already done; the aggregate appears in the Observatório de IA do CIn, with 2,943 responses in the September 14, 2026 extract. The survey is cross-sectional — its cycles do not follow the same people, and a curve that rises from one cycle to the next does not demonstrate progress caused by any training.

04 · Fluency

Where the scale stops

The measure does not assess the quality of what AI produces, no instrument sends conversation content to any server, and no individual result is disclosed: on public dashboards, cells with fewer than five responses are suppressed, which in the Observatory left 1,394,729 cells out of the totals. An unavailable combination does not mean zero.

Two objections are published alongside the index. Goodhart's law: an indicator that becomes a target stops measuring what it represented, and a fluency index is vulnerable through three doors — rewarding volume of use, trivial to manipulate; being read as a seal of approval; and scoring delegation in a way that rewards whoever delegates most outside their own competence. The second: an announced measure functions as a public commitment and can be held accountable for later, including legally.

Competence distillation — delegating code generation charges a price in skill formation, and charges more of those who are just starting out — is a recurring hypothesis in Cabral's writing and has not yet been tested. Testing it requires following the same people over time, and FAROL's current instruments take cross-sectional measurements. Açude, a common-pool-resource game launched in 2026, records the decision round by round; on September 18, 2026, the match still returned a 404.

Frequently asked questions

What is AI fluency?

Artificial intelligence fluency is the speed at which a person or organization moves in its use of AI, measured axis by axis, rather than the level reached on a scale. As with a language, it is an ongoing state: those who stop practicing regress. It is the focus of FAROL, a group of two professors at CIn/UFPE — Giordano Cabral and Filipe Calegário — that has measured fluency along six axes since 2025, on a scale from N0 to N5.

What is the difference between AI fluency and AI maturity?

In FAROL's definition, maturity measures position and fluency measures movement. A maturity model places the organization on a rung of a ladder inherited from 1990s software, built for stable processes. When the state of the art changes every few months, the reference point shifts after the measurement is taken. Fluency measures how fast the organization is moving, and along which axis.

What does it mean to measure through traces rather than self-reports?

A self-reported measure is what a person says about themselves in a questionnaire. A trace-based measure is the record of what they did: a decision within a game, the time spent on each screening item, the tool used in a task. Self-reports are systematically inaccurate — in METR's July 2025 trial, experienced professionals took 19% longer with AI and estimated that they had become 20% faster. The pair appears in three places within the same work: the game, screening telemetry, and observed use.

Does FAROL's fluency index measure the quality of what AI produces?

No. The instruments also do not send conversation content to any server or disclose individual results: in public dashboards, any cell with fewer than five responses is suppressed. Published alongside the index are the two objections the group has not resolved — Goodhart's law and the fact that an announced measure is a promise.

Where can an organization measure its own fluency today?

Through FAROL's public diagnostic assessment, at farol-ia.org/diagnostico. A leader responds on behalf of the organization; it takes three to eight minutes and requires no prior registration. Results are available immediately, with a radar chart of the six axes and an indication of which side of the boundary between generative AI and agentic AI the company is on. It is free, with no consulting engagement involved, and the question bank behind it contains 264 items.