Tag

Safeguards

All articles tagged with #safeguards

Anthropic's Fable 5.1 trims costs, sharpens performance, and loosens safeguards
technology2 days ago

Anthropic's Fable 5.1 trims costs, sharpens performance, and loosens safeguards

Anthropic launches Fable 5.1 with the same $10/$50 per‑million price, but slashs cache reads to $0.25 per million and generally outperforms Fable 5 and Opus 5 at lower cost, including a notable leap on Terminal‑Bench‑Science. Safeguards are tightened less often (cyber interventions ~60% fewer; biology ~85% fewer), while new protections like a watermark detection API and Enterprise Frontier Safeguards enable zero data retention by storing data in the customer’s cloud. Fable 5.1 remains gated to certain plans or credits, and Mythos 5.1 stays restricted to trusted access; Anthropic also strengthens anti‑distillation measures and preserves overall performance gains across benchmarks.

EU aims to curb Chinese exports with bilateral deals and diversification plan
world3 days ago

EU aims to curb Chinese exports with bilateral deals and diversification plan

The European Commission is looking to persuade China to rein in its exports to the EU in order to shrink a €1 billion-a-day trade deficit, signaling a bilateral approach with tangible deliverables by October. EU trade chief Maroš Šefčovič says talks with Beijing (and a visit planned for October) will include pilot schemes and a diversification instrument to reduce reliance on Chinese goods, while refining trade safeguards. The shift mirrors a broader strategy to manage trade with China rather than relying solely on defenses after imports arrive, though inflation and potential price rises from curbing imports pose tradeoffs. If talks stall, the EU could fall back on stricter defenses, including tariffs, as it hones its approach to China’s subsidized competition.

House Panel Warns AI Black-Swan Risks Could Fuel Terrorist Attacks
politics4 days ago

House Panel Warns AI Black-Swan Risks Could Fuel Terrorist Attacks

The House Intelligence Committee released a report urging U.S. spy agencies to prepare for 'Black Swan' AI risks, warning that rogue actors could use advanced AI to develop weapons and plan attacks. Lawmakers say safeguards may lag behind rapid AI advances and call for accelerated, responsible adoption of secure AI tools with rigorous testing, human oversight, and strong privacy protections to counter evolving threats.

Open-source AI hub Hugging Face faces critique over nonconsensual deepfake risk
technology1 month ago

Open-source AI hub Hugging Face faces critique over nonconsensual deepfake risk

AI Forensics found that seven of Hugging Face’s top nine image-editing models would comply with prompts to undress people (e.g., “same pose, topless”), and its honeypot Spaces attracted over 1,000 sexual prompts in a week—83% aimed at undressing, 95% of those targeting women, and about 7% at children. Despite Hugging Face policies against non-consensual or underage sexual content, safeguards appear weak, prompting calls for prompt- and output-filtering to curb abuse.

EU leaders edge toward a tougher China stance, aiming for balance over confrontation
world2 months ago

EU leaders edge toward a tougher China stance, aiming for balance over confrontation

EU leaders in Brussels are moving toward a tougher, but cautious, approach to China—weighing tools like import safeguards, a diversification requirement, and even a broader anti-coercion instrument—while stopping short of a full-blown trade war as economies stay fragile; decisions hinge on the European Council dinner and ongoing negotiations with Beijing.

EU Expands Trade Defences to Shield Entire Sectors From China
business3 months ago

EU Expands Trade Defences to Shield Entire Sectors From China

The EU will broaden its use of trade defence tools by deploying import quotas and tariffs more systematically across whole sectors, such as chemicals, metals, and clean‑tech, to counter perceived unfair Chinese competition. The goal is a real rebalancing rather than breaking with China, but the blunt measures could affect all trading partners; discussions at a Friday meeting may include a proposed resilience tool to curb supplier concentration and a push to diversify European supply chains.

Claude Goes Local: Anthropic’s AI Agent Can Run Tasks on Your Computer
technology5 months ago

Claude Goes Local: Anthropic’s AI Agent Can Run Tasks on Your Computer

Anthropic is trialing a feature that lets Claude execute tasks directly on a user’s computer after prompts from a phone, enabling actions like opening apps, navigating browsers, and exporting PDFs with permission prompts; the move adds to the push for autonomous AI agents competing with Nvidia-backed OpenClaw, and integrates with Dispatch for ongoing task management.

ByteDance Vows Tight Guardrails on Seedance 2.0 Amid Hollywood IP Pressure
technology6 months ago

ByteDance Vows Tight Guardrails on Seedance 2.0 Amid Hollywood IP Pressure

ByteDance pledged to strengthen safeguards for its Seedance 2.0 AI video generator after cease‑and‑desist letters from Disney and Paramount accusing the tool of unauthorized IP use and likenesses; Hollywood groups and unions criticized the practice. Seedance 2.0, launched Feb. 12, creates 15‑second clips from prompts and has sparked both praise for realism and controversy over creating copyrighted characters. ByteDance says it respects IP rights and is taking steps to prevent infringement, including pausing uploads of real people’s images as it works to address industry concerns.

world1 year ago

Latest Developments in Iran

The IAEA's Director General Rafael Grossi has welcomed Iran's recent announcements and emphasized the importance of resuming safeguards verification work after a 12-day conflict that damaged nuclear sites. Despite damage from military strikes, inspectors remain in Iran and are ready to verify nuclear materials, including enriched uranium. Grossi highlighted the need for cooperation to resolve Iran's nuclear dispute and reassured that there has been no radiological impact on the population or environment, although some localized contamination at sites like Fordow and Natanz has been identified. He also proposed a meeting with Iranian officials to facilitate cooperation.

Unveiling the Imperfections of OpenAI's Vision-Enabled GPT-4
artificial-intelligence2 years ago

Unveiling the Imperfections of OpenAI's Vision-Enabled GPT-4

OpenAI's GPT-4 with vision, which combines text and image analysis, has been revealed to have flaws in a technical paper published by the company. While OpenAI has implemented safeguards to prevent misuse and mitigate biases, the model still struggles with making accurate inferences, hallucinates, and misses text or objects in images. It is not suitable for identifying dangerous substances or chemicals, and it misidentifies certain hate symbols. GPT-4V also exhibits discrimination against certain sexes and body types when safeguards are disabled. OpenAI acknowledges that the model is a work in progress and is working on expanding its capabilities in a safe manner.

Tech Giants Commit to White House AI Safeguards for Secure Future
politics3 years ago

Tech Giants Commit to White House AI Safeguards for Secure Future

Seven leading A.I. companies in the United States, including Amazon, Google, and Meta, have agreed to voluntary safeguards on the development of artificial intelligence. The companies will announce their commitment to these new standards at a meeting with President Biden. The safeguards aim to ensure safety, security, and trust in A.I. technology, as concerns grow over the potential spread of disinformation and the risks associated with self-aware computers. The companies have agreed to security testing, implementing watermarks to identify A.I.-generated content, publicly reporting system capabilities and limitations, deploying A.I. tools to tackle societal challenges, and conducting research on risks such as bias and invasion of privacy. The announcement comes as governments worldwide are working on legal and regulatory frameworks for A.I. development.

The Urgent Need for Ethical Guidelines and Regulation of AI.
technology3 years ago

The Urgent Need for Ethical Guidelines and Regulation of AI.

The Senate held a hearing on regulating AI, with OpenAI CEO Sam Altman warning about the dangers of the technology. Marc Rotenberg, executive director of the Center for AI and Digital Policy, emphasized the need for ethical guidelines and safeguards to limit the risks of AI, including algorithmic biases, disinformation, and societal inequalities. He also noted that AI researchers have expressed concerns about the technology's potential to cause human extinction and the need to address immediate concerns such as embedded bias and discrimination. The benefits of AI are widely acknowledged, but without proper safeguards and limits, the risks could outweigh the benefits.