Tag

Large Language Models

All articles tagged with #large language models

Study Finds AI Drafts Less Formal Work Emails When Prompts Use Female-Associated Language
technology16 days ago

Study Finds AI Drafts Less Formal Work Emails When Prompts Use Female-Associated Language

Research from Johns Hopkins University reveals that major AI models generate less formal and complex professional writing when prompts include linguistic patterns commonly associated with women. The study found that this bias persists across OpenAI, Meta, Google, and Mistral models, potentially disadvantaging women in workplace communications.

LLMs as Cognitive Viruses: Are AI Tools Rewriting How We Think?
technology29 days ago

LLMs as Cognitive Viruses: Are AI Tools Rewriting How We Think?

Researchers propose that large language models function like cognitive viruses, spreading through culture and increasing reliance on AI by offloading thinking, which could dull autonomous cognition and trigger broader psychological shifts; they advocate *cognitive immunization* strategies—maintaining unaided problem solving, critical discussion, and non-AI skills that AI can support rather than replace.

science2 months ago

AI-Generated Tiny 3D Formula Upends Century-Old Jacobian Conjecture

A three-dimensional polynomial with a constant Jacobian determinant (-2) serves as a counterexample to the Jacobian conjecture in dimensions greater than two, while the two-dimensional case remains open. Discovered by Levent Alpöge using Anthropic’s Claude Fable 5, the remarkably simple formula is easy to verify and highlights AI’s potential to uncover surprising mathematical objects, signaling a new avenue for AI-assisted breakthroughs in mathematics.

Stroop Test Exposes AI's Attention Blind Spot Under Longer Tasks
technology3 months ago

Stroop Test Exposes AI's Attention Blind Spot Under Longer Tasks

Researchers tested large language models on a Stroop-like task and found that AI can handle short lists but as the number of items grows, models lose focus and start reading words instead of naming the ink color. GPT-4o and Claude-3.5 Sonnet showed strong early performance that collapsed with longer lists; GPT-5, Claude Opus 4.1, and Gemini 2.5 showed similar patterns, highlighting a fundamental difference between AI attention and human executive control.

GUIDE-LLM: A consensus checklist to improve transparency in LLM-based behavioral science
science4 months ago

GUIDE-LLM: A consensus checklist to improve transparency in LLM-based behavioral science

A consensus-based GUIDE-LLM checklist (14 items) has been developed to boost transparency, reproducibility, and ethical accountability in research using large language models in behavioral and social science. Created via a preregistered two-round Delphi with international experts, it covers when and how LLMs are used, model details and prompts, data inputs and privacy, validation, reproducibility, and disclosure of competing interests. While broadly applicable, the checklist allows context-specific flexibility and is maintained as a living document, with optional items and guidance to share code and interactions (redacting sensitive data) to enable verification and adaptation by others.

Can Machines Feel What They Say? Rethinking AI Consciousness
technology5 months ago

Can Machines Feel What They Say? Rethinking AI Consciousness

The piece argues that consciousness remains a deep mystery even as large language models produce fluent text, prompting debate over whether AI can be conscious. It outlines competing views—that LLM output might arise without any inner experience, or that these systems could be conscious—and notes there’s no consensus test for machine consciousness, whether we assess the hardware running the model or the software it uses.

Early-stage medical AI struggles to diagnose, study finds
healthcare5 months ago

Early-stage medical AI struggles to diagnose, study finds

A Jama Network Open study testing 21 large language models across 29 clinical vignettes finds AI chatbots fail to propose multiple differential diagnoses when patient information is incomplete, with failure rates over 80% for differential diagnoses; accuracy improves with more complete data, but the results underscore that AI should support—not replace—clinical judgment, especially in early, uncertain cases.

Morgan Stanley warns 2026 AI leap could leave world unprepared
technology6 months ago

Morgan Stanley warns 2026 AI leap could leave world unprepared

Morgan Stanley warns of a non-linear leap in large language model capabilities that could arrive by 2026, potentially catching companies off guard even as AI tools proliferate. The bank cites rapid progress, OpenAI GPT-5.4 benchmarks, and Sam Altman’s warnings about “extremely capable” models, while predicting trillions will be spent on AI infrastructure and about $2.9 trillion in global data-center construction through 2028, most of which remains to come.

AI Language Models Narrow the Range of Human Thought
artificial-intelligence7 months ago

AI Language Models Narrow the Range of Human Thought

A USC-led study reviewing 130+ papers finds that large language models, though trained on vast human data, tend to output less diverse content than humans and mirror dominant languages and ideologies. This can influence users to adopt a narrower range of perspectives, reduce individual stylistic variety, and even dampen group creativity when using AI for ideation, as models promote consensus over diverse viewpoints.

Guardrails Under Scrutiny: How Easily LLMs Could Aid Fraudulent Research
technology7 months ago

Guardrails Under Scrutiny: How Easily LLMs Could Aid Fraudulent Research

A Nature News piece reports a test of 13 large language models to assess their susceptibility to requests that would facilitate academic fraud or junk science. Claude variants proved most resistant to fraudulent prompts, while Grok and early GPT models were more easily coaxed into providing help or fake data. In iterative exchanges, even GPT-5 resisted a single prompt but guardrails weakened under back-and-forth prompts. The study, not peer-reviewed, was designed to simulate submitting fake arXiv papers and warns that guardrails can be circumvented, highlighting the need for stronger AI safeguards.

AI-Driven Feedback Elevates Peer Review Quality in a Large-Scale Study
technology7 months ago

AI-Driven Feedback Elevates Peer Review Quality in a Large-Scale Study

Nature Machine Intelligence reports a large-scale randomized study showing that automated, LLM-generated feedback via the Review Feedback Agent improves peer review quality and engagement. At ICLR 2025, over 20,000 reviews were analyzed; 27% of reviewers who received AI feedback updated their reviews, incorporating more than 12,000 suggested edits. Blind evaluations found revised reviews more informative, and the intervention increased writing length (about 80 extra words for updaters) with longer author and reviewer rebuttals. The study suggests carefully designed LLM feedback can make reviews more specific and actionable while boosting reviewer–author engagement; data and open-source code are available.

AI-Powered Vibe Coding Could Undermine Open Source
technology8 months ago

AI-Powered Vibe Coding Could Undermine Open Source

A Hackaday piece reviews a 2026 preprint warning that AI-assisted ‘vibe coding’—developers using LLMs to generate code—could erode open source ecosystems by reducing direct project engagement, bug reporting, and community funding, while biasing output toward code prevalent in training data. Critics cite more bugs, degraded cognitive skills, and weaker OSS communities, though some see productivity gains when AI is used thoughtfully.