Study compares attention in 13 LLMs
Published on October 12, 2025, the report compares attention and linear DeltaNet attention proportions in 13 LLMs; two of 12 layers using attention, or 17%, produced the best result in the experiments.
Source: Estudo compara níveis de attention em 13 LLMs (x.com). Text prepared with AI from this source.
What happened and what to do
On October 12, 2025, a study report said its author trained 13 LLMs with different proportions of attention and linear DeltaNet attention. In the described experiments, the configuration with 17% attention—two of 12 layers—had the best result. The report does not specify metrics, tasks, or further model details. To consult and verify the information, look up the original record under the title “Estudo compara níveis de attention em 13 LLMs” and check the date and results presented there.
For a team developing models, this case points to the value of testing architecture configurations under controlled conditions, without assuming the result applies to other tasks. A company can implement an experimentation workflow that records configurations, data, metrics, and costs, compares alternatives, and displays results in a dashboard to inform engineering decisions.
How the consultancy can help
Wendelmaques can assess a company’s current model evaluation process, define an experimentation scope, and implement test pipelines, metric logging, and dashboards. It can also support operation of this process on the company’s infrastructure.
Next step
Send a short description of the model or process your team wants to evaluate. Wendelmaques can review the context and prepare a proposal scoped for diagnosis and implementation.
Consulting for your project
Infrastructure review, deployment and ongoing operations, with scope and pricing defined in the proposal.
Quoted per project
Request a proposal