AI agents are increasingly auditing science literature, exposing decades-old errors in a long-trusted chemistry reference database and in conference papers. While AI can speed up error detection and reproducibility checks, human oversight remains essential due to occasional mistakes and misreadings.
Nature reports that PubPeer is launching a project to post replication studies on its platform and link them to the original papers, in an effort to make replication efforts more visible and speed scientific self-correction. The plan involves about 2,400 replication studies added in batches of ~100 via the FORRT Library of Reproduction and Replication Attempts (FLoRA), highlighting both successful and failed replications. Authors of the papers are being notified, with mixed expectations about PubPeer comments, but supporters argue that publicly confirming a study’s validity can benefit the research community by reducing wasted effort and improving credibility.
The article shows how researchers use everyday kitchen items and simple gear to make field science more robust, reproducible, and accessible: a soup ladle on a pole and a strainer to collect and clean brine samples; a jewellery chain to estimate soil roughness; and kite-based surveys as durable, low-cost alternatives to drones. It emphasizes improvisation in remote work, contrasts high-tech and low-tech methods, and highlights global collaborations (like CrustNet) built on shared, widely available protocols to democratize scientific data collection across diverse sites.
A consensus-based GUIDE-LLM checklist (14 items) has been developed to boost transparency, reproducibility, and ethical accountability in research using large language models in behavioral and social science. Created via a preregistered two-round Delphi with international experts, it covers when and how LLMs are used, model details and prompts, data inputs and privacy, validation, reproducibility, and disclosure of competing interests. While broadly applicable, the checklist allows context-specific flexibility and is maintained as a living document, with optional items and guidance to share code and interactions (redacting sensitive data) to enable verification and adaptation by others.
Large, cross-lab projects in infant- and animal-cognition are being used to tackle the psychology reproducibility crisis, with initiatives like ManyBabies, ManyDogs, and ManyBirds increasing statistical power and diversity, though results have been mixed and sometimes contradict earlier findings.
Eighteen teams analyzed the same Neuropixels dataset and largely disagreed on ripple density across brain areas, despite using defensible methods. The divergence arose from differences in how concepts were defined, which algorithms were used, and the parameters chosen, revealing substantial analytical variability. The effort spurs the CON²PHYS project to quantify conceptual disagreement and push for transparency, reference pipelines, and reporting standards to ensure conclusions are robust to analytical choices.
An NIST redo of the 2007 BIPM measurement of the gravitational constant G, using a blinded-envelope approach to avoid bias, yields a result close to the French value but with a 0.0235% discrepancy after adjustments; Schlamminger also identifies a newly observed spurious torque driven by temperature gradients and residual gas in the vacuum, suggesting unaccounted biases in the uncertainty budget and underscoring the ongoing challenge of precisely measuring G and the importance of reproducibility.
An opinion piece cautions that rapid, uncritical adoption of AI and large language models in science is boosting output while narrowing inquiry, risking lower-quality results and erosion of tacit training for early-career researchers. It calls for guardrails to preserve hands-on apprenticeship, ensure responsible oversight of AI-assisted workflows, and use metrics that reflect true scientific understanding rather than sheer productivity.
A genome-wide check of 611 samples from 341 mouse strains in the NIH/MMRRC network found that 47% did not match their reported identities, exposing widespread mislabeling and genetic drift that could undermine the reproducibility of studies relying on these models.
A Europe-backed NanoBubbles project is funding nanoscientists to replicate a 2012 study that carbon quantum dots can sense copper ions inside living cells, the first large-scale replication effort in the physical sciences aimed at the reproducibility crisis; initial attempts failed to reproduce the reported fluorescence change, illustrating how small impurities, incomplete protocols, and cross-lab variation can affect results, as the ERC-backed effort seeks self-correction in science.
Chemical engineers defend the importance of stirring in chemical reactions, arguing that while some small-scale, homogeneous reactions may not require mixing, it remains critical for reproducibility, safety, and scalability in industrial and heterogeneous systems, especially to prevent hazards like hotspots and runaway reactions. The debate was sparked by a study claiming stirring is unnecessary for certain organic reactions, but experts emphasize that mixing is essential in many practical scenarios, particularly at larger scales.
A collection of articles and books explore the integration of artificial intelligence (AI) into scientific research, discussing its potential impact on various disciplines, the ethical implications, and the challenges related to reproducibility and interpretability. The use of large language models in research is critiqued, with attention to issues such as bias, distortions of human beliefs, and limitations in predicting scientific replicability. Additionally, the application of AI in literature reviews, protein structure prediction, and other scientific domains is examined, highlighting both the opportunities and the need for careful consideration of the implications of AI in scientific discovery.
A study involving over 200 biologists analyzing the same ecological data set has revealed significant variations in their results, highlighting the impact of scientists' analytical choices on research outcomes. The findings emphasize the need to avoid relying solely on individual studies and results, as they may not provide a comprehensive understanding of a particular phenomenon. The study's authors suggest that transparency regarding analytical decisions and conducting robustness tests could help address the issue of reproducibility in ecology.
A study in the field of ecology has found empirical evidence of widespread exaggeration bias and selective reporting, highlighting concerns about the reproducibility of research findings in the discipline. The study examined the prevalence of these biases in ecological research and their potential impact on effect sizes, statistical power, and the occurrence of type M (magnitude) and type S (sign) errors. The findings suggest that publication bias and the pressure to report statistically significant results may contribute to the exaggeration of effect sizes and the suppression of non-significant findings. The study emphasizes the need for transparency, reproducibility, and improved statistical reporting practices in ecology to ensure the credibility and reliability of research findings.
Data scientists play multiple roles in collaborations, including data analysis, data acquisition, software development, and project management. However, misunderstandings and undervaluing their contributions can hinder effective collaboration. To improve working relationships, it is important to establish a communication plan, communicate openly, learn each other's jargon, encourage questions, and use creative communication methods. Additionally, setting a timeline, avoiding scope creep, planning for data storage and distribution, prioritizing reproducibility, documenting everything, and developing a publishing plan are crucial. Embracing creativity, sharing knowledge, and recognizing when a project has run its course are also important for successful interdisciplinary collaborations in data science.