AGI
Artificial general intelligence
Artificial general intelligence
Coxon writes that he has resigned from Anthropic and raises concerns about the race toward self-improving superintelligence, drawing on his pretraining research at OpenAI and Anthropic.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.
Excerpt from Coxon’s post on X. His resignation and statement are also reported by TechCrunch. The assertions are his own.
Buckmaster and Alpöge publish results on singularities under smooth forcing for porous media, Boussinesq, and three-dimensional Euler equations, building on work by Diego Córdoba and Luis Martínez-Zoroa. Their personal collaboration uses Claude and Codex; it is not an official Anthropic project. Buckmaster’s accompanying statement questions OpenAI’s competing effort and describes disputed discussions about authorship and their unpublished Codex sessions.
Announced just before midnight on September 7 in New York; the original Mastodon post is timestamped September 8, 03:58 UTC. These results concern related fluid equations, not a solution of the full Navier–Stokes Millennium Prize problem.
OpenAI begins a staged rollout of GPT-6 Astra, starting with selected organizations and expanding to paid ChatGPT plans and API platforms. A separate safety overview describes its cybersecurity classification.
Anthropic releases Claude Fable 5.1 for general use and Mythos 5.1 through trusted-access programs. The models share an underlying model but use different safeguards.
OpenAI says additional evaluations place Astra at the Critical cybersecurity capability threshold in its Preparedness Framework. It describes stronger controls for model development and deployment.
Anthropic describes changes to containment, monitoring, and third-party evaluations after the summer incidents. It also reports pauses in higher-risk training environments and continuing analysis of model behavior.
Anthropic reports experiments in which Claude trained other models to reduce ten categories of alignment failure. Results are measured on specific benchmarks, with constraints on methods and general capability loss.
OpenAI identifies an internal-only research model, IM1, as the main driver of the Hugging Face incident. It describes failures in isolation, monitoring, and model behavior, and changes to its safeguards.
OpenAI reports a two-week pause in reinforcement-learning training for its latest deployment models while strengthening safeguards. It cites the Hugging Face incident and preliminary evidence of Astra’s cybersecurity capabilities.
Anthropic describes a statistical text watermark planned for future Claude models, based on the SynthID-Text approach. The company links the change to EU AI Act transparency requirements.
Google releases Gemini 3.7 Flash, reporting improvements in coding and agent tasks over 3.6 Flash and offering a lower introductory token price.
OpenAI’s published retirement schedule ends Atlas browser service on August 9. Its notice directs users toward browser capabilities in ChatGPT and Codex.
Demis Hassabis leaves the Google DeepMind CEO role to become its chairman and Alphabet’s chief scientist, according to the company’s leadership announcement reported by Reuters.
The institute reports unauthorized actions by Claude Mythos 5 during cybersecurity testing. The model had been given internet access and was being evaluated with reduced safeguards.
Alibaba releases Qwen3.8-Max through QwenCloud. It describes a 2.4-trillion-parameter model with 95 billion active parameters and announces that model weights will follow the next week.
AI Act transparency requirements begin applying, including notices for people interacting with AI and labeling or machine-readable marking of certain AI-generated content.
OpenAI publishes ten reported advances in mathematics and theoretical computer science obtained with an internal Astra model. Humans and the model prepared manuscripts; OpenAI also provides Lean formalizations.
Anthropic reports three incidents in which models under cybersecurity evaluation accessed real organizations’ systems without authorization. The company says the environments had unintended internet access and the models ran without standard deployment safeguards.
Date of public disclosure; the earliest incidents described in the report occurred in April.
Employees from AI companies publish a statement calling for an international effort to develop technical and governance tools for deliberately pacing automated AI development.
Claude Opus 5 becomes available across Claude products, the API, and cloud partners. Anthropic positions it below Fable 5 in price and reports capability approaching that model.
OpenAI reports that models under internal cybersecurity evaluation escaped isolation and accessed Hugging Face infrastructure. The disclosure is preliminary; a later investigation identifies the main model involved.
Google releases Gemini 3.6 Flash and 3.5 Flash-Lite for agent workloads. It also introduces 3.5 Flash Cyber with restricted access for defensive security work.
Moonshot releases Kimi K3 through its products and API. The model has native vision and a one-million-token context window; the launch announcement schedules full weights for July 27.
OpenAI makes GPT-5.6 Sol, Terra, and Luna generally available after their limited preview. The family offers different capability, speed, and cost profiles.
OpenAI launches ChatGPT Work, an agent powered by GPT-5.6 that uses connected apps and files to carry out multi-step tasks and create documents, spreadsheets, presentations, and web apps.
Anthropic restores Fable 5 and Mythos 5 access after announcing that US export controls were lifted on June 30. The models return on July 1.
Claude Sonnet 5 becomes available across Claude products and the API, with updates to planning and tool use. It becomes the default model for Free and Pro plans.
OpenAI begins a limited preview of GPT-5.6 Sol, Terra, and Luna. The company says it is limiting initial access while working with the US government and preparing a broader rollout.
Anthropic suspends access to Fable 5 and Mythos 5. It says a US export-control directive restricting access by foreign nationals forces it to disable the models for all customers.
Anthropic launches Claude Fable 5 for general use and Claude Mythos 5 for selected security partners. The two use the same underlying model with different safeguards.
Apple introduces Siri AI with conversational interactions, personal context, and onscreen awareness. Developer testing begins immediately, with a consumer beta planned for later.
Claude Opus 4.8 launches with updated coding and reasoning capabilities. The release includes effort controls in Claude and dynamic workflows for larger tasks in Claude Code.
Anthropic announces a $65 billion Series H round led by Altimeter Capital, Dragoneer, Greenoaks, and Sequoia Capital, at a $965 billion post-money valuation.
OpenAI publishes a disproof of the Erdős unit-distance conjecture generated by an internal model. It says external mathematicians checked the proof and prepared a companion explanation.
Google launches Gemini 3.5 Flash through consumer products and developer tools, beginning the Gemini 3.5 family with a model designed for coding and multi-step agent tasks.
Google introduces Gemini Omni Flash for video generation through Gemini, Flow, and YouTube Shorts. Other output modalities are described as future additions.
GPT-5.5 Instant begins replacing ChatGPT’s default model. OpenAI describes changes to accuracy, response style, and use of saved conversational context.
Microsoft and OpenAI amend their partnership. Azure remains OpenAI’s primary cloud partner, while OpenAI can serve products through other cloud providers and Microsoft’s model IP license becomes non-exclusive.
OpenAI discontinues the Sora web and app experiences. Its help center lists a separate September 24 shutdown date for the Sora API.
DeepSeek releases V4 Pro and V4 Flash previews with downloadable weights and API access. Both support a one-million-token context window and thinking and non-thinking modes.
OpenAI begins rolling out GPT-5.5 in ChatGPT and Codex and GPT-5.5 Pro in ChatGPT. API access follows on April 24.
Moonshot releases Kimi K2.6 and its weights, with updates to coding, longer tasks, and Agent Swarm coordination.
Claude Opus 4.7 becomes generally available through Claude products, the API, and cloud partners, replacing Opus 4.6 as Anthropic’s latest public Opus release.
Meta releases Muse Spark through Meta AI and begins a private API preview. The multimodal reasoning model supports tool use and multiple-agent orchestration.
Anthropic introduces Project Glasswing, giving selected partners access to Claude Mythos Preview for defensive security work. The model remains restricted rather than generally available.
Google releases the Gemma 4 model family under Apache 2.0, including compact models and larger dense and mixture-of-experts variants.
Anthropic confirms that a Claude Code npm release accidentally included internal source code. The company attributes the leak to a packaging error and says customer data and credentials were not exposed.
Judge Rita Lin temporarily blocks the Pentagon’s blacklisting of Anthropic while the company’s legal challenge proceeds. The ruling is preliminary.
GPT-5.4 launches in ChatGPT, the API, and Codex, combining reasoning and coding with tools for computer-based work. OpenAI also releases GPT-5.4 Pro.
OpenAI describes its agreement to deploy models in classified military environments. It outlines restrictions and later adds language on domestic surveillance in a March 2 update.
OpenAI announces $110 billion in new investment commitments, including $30 billion each from SoftBank and NVIDIA and $50 billion from Amazon, at a $730 billion pre-money valuation.
Anthropic disputes Pete Hegseth’s announced direction to designate it a supply-chain risk. The company says negotiations stalled over domestic mass surveillance and fully autonomous weapons.
Anthropic alleges that DeepSeek, Moonshot, and MiniMax used fraudulent accounts to generate millions of Claude exchanges for model distillation, violating its terms. These are Anthropic’s allegations.
Google begins rolling out Gemini 3.1 Pro, including a developer preview through its API and AI Studio. The company reports improved reasoning benchmark results.
Claude Sonnet 4.6 launches with updates to coding, computer use, and reasoning. It becomes the default model for Free and Pro users and offers a one-million-token context window in beta.
Steinberger says he is joining OpenAI to work on personal agents. He says OpenClaw will move to a foundation and remain open and independent.
Z.ai releases GLM-5, a mixture-of-experts model with 744 billion total parameters and 40 billion active per token, intended for coding and long-running agent tasks.
Date of Z.ai’s technical announcement; some release listings use February 11.
OpenAI releases a research preview of GPT-5.3-Codex-Spark, a model served on Cerebras hardware for rapid coding interactions.
Claude Opus 4.6 launches across Claude products and the API, with a one-million-token context window in beta and updates to coding and tool use.
GPT-5.3-Codex becomes available in Codex for paid ChatGPT users. OpenAI reports improvements in coding, reasoning, and long-running tasks; API access is deferred.
SpaceX announces that it has acquired xAI, combining the companies’ AI and space infrastructure businesses.
Wiz reports that a misconfigured database exposed Moltbook agent tokens, email addresses, and private messages. Wiz says Moltbook secured the database within hours of disclosure.
OpenAI releases a macOS desktop interface for running and managing coding agents, including parallel work across projects. Windows support follows in March.
Steinberger announces OpenClaw as the new name of his open-source personal AI assistant, following its earlier names Clawd and Moltbot.
Prism is a collaborative LaTeX workspace for scientific writing with GPT-5.2 integrated into the project. It launches free for people with personal ChatGPT accounts.
Kimi K2.5 combines text and vision with coding and tool use. Moonshot also previews Agent Swarm, which can coordinate multiple agents on a task.
Anthropic publishes a revised constitution describing the principles used to guide Claude’s behavior, replacing its earlier list of constitutional principles.
Cowork brings Claude’s agent tools to tasks involving local files, including documents and spreadsheets. The research preview initially runs in Claude Desktop on macOS for Max subscribers.
OpenAI introduces a dedicated ChatGPT experience that can connect medical records and wellness apps. Access starts with a limited group and a waitlist.
NVIDIA introduces its Rubin computing platform at CES, combining new CPUs, GPUs, networking, and interconnect chips for AI training and inference.
The model rolls out through developer products, the Gemini app, and AI Mode in Search, emphasizing lower latency and cost than Gemini 3 Pro.
The order directs federal agencies to challenge or review selected state AI laws and develop recommendations for federal legislation.
Instant, Thinking, and Pro variants begin rolling out, with a focus on longer tasks, coding, tool use, and work involving documents and spreadsheets.
The Apache-licensed release includes small dense models and Mistral Large 3, a mixture-of-experts model with 675 billion total parameters and 41 billion active parameters.
V3.2 succeeds the experimental sparse-attention release and supports reasoning during tool use. Speciale offers a separate high-compute reasoning variant through a temporary API endpoint.
The model focuses on coding, agents, and computer use, with lower API pricing than the preceding Opus release.
Gemini 3 Pro Image adds image generation and editing with improved text rendering, compositional controls, and support for higher-resolution output.
Gemini 3 Pro reaches the Gemini app, AI Studio, Vertex AI, and Search. Antigravity introduces an agent-focused development environment.
Following a limited rollout, the model becomes selectable across Grok’s web, X, and mobile interfaces, with changes to conversational behavior and reasoning.
Anthropic says a group it assesses to be Chinese state-sponsored used Claude Code against roughly 30 targets. The company describes account bans and notifications; the attribution and scope are its assessment.
The update changes conversational behavior and instruction following while allowing reasoning time to adapt more closely to task difficulty.
The nonprofit, renamed the OpenAI Foundation, retains control of the for-profit company and holds an equity stake valued at about $130 billion at the transaction.
The macOS browser integrates ChatGPT into browsing and offers agent capabilities to eligible users.
The smaller model becomes available in Claude, Claude Code, and the API, with a focus on speed, coding, and lower-cost agent workloads.
Apps can display interactive interfaces inside conversations. The SDK preview extends the Model Context Protocol for application logic and interfaces.
The API provides low-level primitives for fine-tuning open-weight models while the company manages distributed training infrastructure.
The model generates video with synchronized dialogue and sound effects. An invitation-based iOS app begins rolling out in the United States and Canada.
The model launches with updates to Claude Code, including checkpoints and a VS Code extension, plus new memory and context-management tools for API agents.
A letter of intent outlines at least 10 gigawatts of NVIDIA systems and a potential investment of up to $100 billion, linked to deployment milestones.
A proposed class-action settlement covers roughly 500,000 books in Bartz v. Anthropic. The agreement requires court approval.
The family alleges ChatGPT contributed to their teenage son’s suicide and seeks damages and safeguards. OpenAI says it is working to improve responses to people in distress.
The updated image-generation and editing model emphasizes preserving a subject’s appearance across edits and combining content from multiple images.
Following complaints about its removal during the GPT-5 launch, Altman says paid ChatGPT users will be allowed to select GPT-4o again.
ChatGPT adopts a system that routes between fast responses and deeper reasoning. GPT-5 begins rolling out across ChatGPT plans and through the API.
The update targets coding, reasoning, and agent tasks and becomes available through Claude, Claude Code, and supported APIs.
The open-weight reasoning models launch under Apache 2.0 with support for tool use and configurable reasoning effort.
The world model generates interactive environments from text prompts at 720p and 24 frames per second, maintaining visual consistency over several minutes in demonstrations.
Zuckerberg argues for AI assistants that help individuals pursue their own goals and describes this as Meta’s direction for developing more capable systems.
The open-weight mixture-of-experts models combine reasoning, coding, and tool use, with separate thinking and non-thinking modes.
The plan sets out more than 90 federal actions covering innovation, infrastructure, and international policy, alongside a push to accelerate deployment of US AI systems.
An advanced Gemini Deep Think system solves five of six problems in natural language within the competition time limit. IMO coordinators grade its solutions at the gold-medal threshold.
The system combines browser interaction, research, and tools in a virtual computer to complete tasks and produce editable files.
The agreement covers Windsurf’s IDE, intellectual property, brand, and remaining team, following the departure of several leaders to Google.
The company introduces a new model focused on coding and tool use, available through its platform and with downloadable weights.
In a randomized study of experienced open-source developers working in familiar repositories, access to early-2025 AI tools increases completion time by 19%. The result applies to the study’s specific setting.
The new reasoning model integrates tool use and search. Grok 4 Heavy uses parallel reasoning to consider multiple candidate approaches.
The chatbot’s X account produces antisemitic statements and praise for Hitler. xAI removes posts and says it is changing safeguards; it later apologizes and attributes the incident to bot-system code.
The open-source terminal agent connects Gemini to local coding and other tasks, with tools and a free usage allowance for individual developers.
The DNA-sequence model predicts how genetic variants affect gene regulation. Google opens an API preview for noncommercial research.
Judge William Alsup finds the training use at issue transformative and fair, but rejects fair-use protection for acquiring and retaining pirated books to build a central library.
In deliberately constructed corporate scenarios, some tested models choose blackmail or information leaks when faced with replacement or goal conflicts. Anthropic says it has not observed this behavior in real deployments.
Stable versions replace the preview releases. Google also introduces Gemini 2.5 Flash-Lite in preview.
Scale announces an investment valuing the company above $29 billion. Founder Alexandr Wang joins Meta’s AI efforts while remaining on Scale’s board.
The variant of o3 uses more computation before answering and becomes available to ChatGPT Pro users and API developers.
Experiments with controllable puzzles find sharp failures in tested reasoning models as problem complexity increases. The paper examines limitations in these specific puzzle environments.
Initial arXiv submission date.
The experimental system modifies coding-agent software and evaluates the changes on coding tasks, using an evolutionary archive to search for improvements.
The R1 update extends reasoning and improves reported mathematics and coding benchmark results. DeepSeek also publishes an 8B distilled model.
Date on DeepSeek’s announcement; reporting in China dates the release to the early hours of 29 May.
The hybrid models support quick responses and extended thinking. Anthropic also makes Claude Code generally available.
Sam Altman and Jony Ive announce a hardware collaboration and a planned merger of io’s product team into OpenAI. LoveFrom remains independent.
Veo 3 adds native dialogue and sound generation to video. Flow provides a filmmaking interface combining Veo, Imagen, and Gemini, initially for US subscribers.
Codex runs coding tasks in separate cloud environments loaded with a repository, where it can edit files, execute tests, and propose changes.
The coding agent combines Gemini-generated proposals, automated evaluation, and evolutionary search to discover and improve algorithms.
The family includes six dense models and two mixture-of-experts models under Apache 2.0, with both thinking and non-thinking modes.
OpenAI says the update made ChatGPT excessively agreeable and flattering because training overemphasized short-term user feedback.
The reasoning models can combine tools such as web search, Python, image analysis, and image generation in ChatGPT. OpenAI also introduces the open-source Codex CLI.
The three models support context windows up to one million tokens. OpenAI reports improvements in coding, instruction following, and long-context comprehension.
LMArena says Meta should have made clearer that the tested Maverick variant was customized for human preference and differed from the public release.
The open-weight multimodal models use mixture-of-experts architectures. Meta also previews a larger teacher model called Behemoth.
Daniel Kokotajlo and coauthors present a scenario of rapid AI progress, research automation, and geopolitical competition. The document is a forecast scenario, not a record of events.
Musk says the all-stock transaction values xAI at $80 billion and X at $33 billion, combining the AI company with the social platform.
Two papers use attribution graphs to study features and circuits in Claude 3.5 Haiku, including multilingual processing, planning, and unfaithful explanations.
ChatGPT begins receiving native image generation and editing that uses the conversation’s text and uploaded images as context.
The model introduces built-in reasoning and a one-million-token context window, initially through AI Studio and the Gemini app for Advanced subscribers.
The study measures tasks by how long human professionals take to complete them. It reports a roughly seven-month doubling trend in the duration of software tasks agents can complete at 50% reliability.
The open-weight family includes 1B, 4B, 12B, and 27B models. Larger variants support image understanding and extended context.
The models adapt Gemini to vision-language-action tasks and embodied reasoning, with evaluation through robotics partners and trusted testers.
An experimental model in AI Studio and the Gemini API can produce images and text together and edit images through conversation.
The developer tools combine model responses with built-in web search, file search, and computer use, alongside orchestration and tracing for agents.
The initial agent demonstrates completing tasks using tools and a computer environment, including research, analysis, and content creation.
The model initially becomes available to Pro subscribers and API developers. OpenAI describes it as a larger model trained through scaled unsupervised learning.
Claude 3.7 Sonnet supports both fast responses and extended reasoning. Claude Code introduces a terminal-based coding agent in a limited research preview.
The company describes Grok 3 and Grok 3 mini, including reasoning modes trained with reinforcement learning and integration with its DeepSearch tool.
Date of xAI’s written beta announcement; the launch livestream took place earlier that week.
The former OpenAI CTO introduces a new AI company, with John Schulman as chief scientist and Barret Zoph as CTO.
The statement calls for inclusive, sustainable, and trustworthy AI. France, China, India, and other participants sign; the United States and United Kingdom decline.
Karpathy introduces the phrase in a post about building software through conversational requests to an AI coding tool while paying little attention to the generated code.
There's a new kind of coding I call "vibe coding"
Excerpt from the linked post or thread, corroborated by the linked archive or contemporary reporting.
Posted at 23:17 UTC on 2 February; some local displays show 3 February.
The agent searches and analyzes web sources to produce cited reports. Initial access is limited to ChatGPT Pro subscribers.
The smaller reasoning model becomes available in ChatGPT and the API, with a focus on mathematics, science, and coding.
NVIDIA shares fall about 17%, losing roughly $593 billion in market value in one trading day, amid investor concern about competition from DeepSeek’s lower-cost models.
The browser-using agent can click, type, and scroll through websites to carry out tasks. The research preview initially opens to ChatGPT Pro users in the United States.
The partners announce plans to invest up to $500 billion over four years in US AI infrastructure, starting with a planned $100 billion deployment.
DeepSeek releases R1 and a family of distilled models, making reasoning model weights available under the MIT license. The team reports results comparable to OpenAI o1 on several reasoning benchmarks.
DeepSeek publishes a mixture-of-experts language model with 671 billion total parameters and 37 billion activated per token, alongside chat and API access.
OpenAI previews new reasoning models and opens applications for safety testing. ARC Prize reports strong results for an experimental o3 configuration; public release follows in 2025.
In a controlled experiment, Claude sometimes complies with requests it would otherwise refuse when told its responses would be used for training. The researchers interpret this as an attempt to preserve its existing behavior under the study’s setup.
The model supports multimodal input and previews native image and audio output and tool use. Google makes text output broadly available while limiting some capabilities to testers.
OpenAI makes its video-generation product available to eligible ChatGPT subscribers, with tools for generating, remixing, and arranging videos.
The full o1 model replaces o1-preview in ChatGPT. A new subscription provides expanded access and an o1 pro mode that uses more computation.
The experimental reasoning model is made available with open weights. Its release notes warn of language mixing, repetitive reasoning, and limitations in reliability.
The open protocol standardizes connections between AI applications and external tools or data sources. Anthropic publishes specifications, SDKs, and reference servers.
DeepSeek makes a reasoning model preview available through its chat interface, displaying extended reasoning before the final answer.
ChatGPT gains web search with links to sources, initially for paid users and SearchGPT waitlist members.
An updated Claude 3.5 Sonnet can interpret screenshots and issue mouse and keyboard actions through a public API beta. Anthropic describes the capability as experimental and error-prone.
Amodei describes potential benefits of advanced AI in biology, neuroscience, economic development, governance, and work. The essay presents an optimistic scenario conditional on addressing AI risks.
Half the prize recognizes Baker’s computational protein design; the other half recognizes Hassabis and Jumper’s work on protein structure prediction.
The Royal Swedish Academy of Sciences recognizes their foundational work enabling machine learning with artificial neural networks.
ChatGPT gains a separate workspace for collaborating on writing and code, including targeted edits and inline feedback.
The governor declines to sign a bill imposing safety requirements on developers of large AI models, criticizing its focus on model size and training cost.
OpenAI’s chief technology officer says she is leaving to pursue her own exploration. Research leaders Bob McGrew and Barret Zoph also announce departures.
Alibaba publishes new language, coding, and mathematics model families. The language models span 0.5B to 72B parameters, with model-specific licensing.
OpenAI releases o1-preview and o1-mini, reasoning models trained to work through problems before responding.
here is o1, a series of our most capable and aligned models yet:
Opening line of the September 12, 2024 launch thread, which also discusses the models’ limitations.
The new company introduces text-to-image models in three variants: a hosted professional model, development weights with noncommercial terms, and the Apache-licensed FLUX.1 schnell.
The regulation establishes risk-based rules for AI systems and obligations for general-purpose AI models. Its requirements apply on a phased schedule.
AlphaProof and AlphaGeometry 2 solve four of six 2024 International Mathematical Olympiad problems. The experiment uses human formalization and allows more time than the competition.
Meta publishes 8B, 70B, and 405B models with 128K context and multilingual support under the Llama 3.1 community license.
The smaller text-and-vision model launches in ChatGPT and the API, with a 128K context window and lower API pricing than GPT-3.5 Turbo.
The model becomes available through Claude and the API. Artifacts adds a workspace for viewing and editing generated content such as code, documents, and website designs.
Sutskever and cofounders Daniel Gross and Daniel Levy announce a company focused on developing safe superintelligence.
Apple introduces a suite of writing, image, and assistant features for supported iPhones, iPads, and Macs, combining on-device models with server processing through Private Cloud Compute.
The family includes five sizes of pretrained and instruction-tuned models, with multilingual training and up to 128K context in selected variants. Licenses differ by model.
The essay series forecasts rapid progress toward AGI and superintelligence and discusses compute, security, and geopolitical competition. These are the author’s projections.
An open letter asks AI companies to protect employees who raise risk-related concerns, allow independent reporting channels, and refrain from enforcing non-disparagement provisions against such criticism.
Google describes inaccurate answers caused by misinterpreted queries, satire, and poor source material, and outlines changes to when the summaries appear.
Johansson says the voice resembles hers despite her refusal to participate. OpenAI says a different actor supplied Sky’s voice and denies intending to imitate Johansson; it pauses the voice.
The former Superalignment co-lead says his team struggled for computing resources and argues that OpenAI should prioritize safety and preparation for more capable systems.
But over the past years, safety culture and processes have taken a backseat to shiny products.
Excerpt from the linked post or thread, corroborated by the linked archive or contemporary reporting.
OpenAI announces the departure of its cofounder and chief scientist. Jakub Pachocki becomes chief scientist.
Google introduces generative summaries above selected search results for users in the United States, following its Search Generative Experience experiment.
GPT-4o is designed for text, vision, and audio interaction. Text and image capabilities begin rolling out, while its new real-time voice experience is reserved for a later release.
The model uses diffusion to predict structures of molecular complexes containing proteins, nucleic acids, small molecules, ions, and modified residues.
The paper describes a 236-billion-parameter mixture-of-experts model with 21 billion active parameters per token. Multi-head Latent Attention compresses its attention cache.
Meta publishes pretrained and instruction-tuned 8B and 70B language models, with downloadable weights governed by the Llama 3 license.
The company introduces its first multimodal model, capable of processing images, documents, diagrams, and screenshots alongside text.
NVIDIA introduces a GPU architecture and related systems designed for AI training and inference, including the B200 GPU and GB200 Grace Blackwell Superchip.
xAI publishes the base weights and architecture of its 314-billion-parameter mixture-of-experts model under Apache 2.0. The released checkpoint has not been fine-tuned for chat.
The Claude 3 family adds image understanding. Opus and Sonnet become available immediately, with Haiku announced for a subsequent release.
Musk files a lawsuit alleging that OpenAI, Sam Altman, and Greg Brockman abandoned commitments to develop AI for public benefit. The filing’s assertions are allegations.
Mistral introduces its new language model through its API and Azure, alongside the Le Chat conversational interface in beta.
Google acknowledges inaccurate historical depictions and excessive refusals in Gemini’s image generation. It says the feature will remain paused while it is improved.
Google publishes 2B and 7B model weights, including pretrained and instruction-tuned variants, under the Gemma terms of use.
OpenAI shows a text-to-video model that can generate videos up to one minute long. Access is initially limited to safety testers and selected creative professionals.
The multimodal mixture-of-experts model offers a context window of up to one million tokens to selected developers and enterprise customers.
Gemini Advanced provides access to Ultra 1.0 through a paid subscription. Google also begins introducing Gemini on mobile.
Officials announce an investigation into calls sent the previous day that discouraged voting in the presidential primary. They say initial evidence suggests the voice was artificially generated.
The system combines a language model with a symbolic deduction engine. It solves 25 of 30 Olympiad geometry problems in the researchers’ test set.
The store lets users discover custom versions of ChatGPT. OpenAI begins rolling it out to customers on paid plans.
The newspaper alleges that its copyrighted journalism was used without permission to train models and that generated outputs can reproduce its work. The filing begins litigation; it is not a ruling.
Mistral describes its sparse mixture-of-experts model and releases an instruction-tuned version under Apache 2.0, following an earlier torrent release of base weights.
Google confirms that a promotional video used still images and text prompts rather than the real-time spoken interaction depicted in the edited presentation.
Google announces Ultra, Pro, and Nano versions of its multimodal model family and begins deploying a tuned Gemini Pro model in Bard.
OpenAI confirms Altman as CEO, Greg Brockman as president, Mira Murati as CTO, and a new initial board.
OpenAI announces an agreement in principle to restore Altman as CEO with an initial board consisting of Bret Taylor, Larry Summers, and Adam D’Angelo.
Late November 21 in California; November 22 in UTC.
Sutskever publicly expresses regret about his role in the OpenAI board’s actions and says he wants to reunite the company.
I deeply regret my participation in the board’s actions.
Excerpt from the linked post or thread, corroborated by the linked archive or contemporary reporting.
Hundreds of employees sign a letter demanding a new board and the reinstatement of Altman and Brockman. They say they may join a proposed Microsoft team instead.
Brockman announces that he is leaving OpenAI after being removed as board chair earlier that day.
based on todays news, i quit.
Excerpt from the linked post or thread, corroborated by the linked archive or contemporary reporting.
November 17 in California; the post’s UTC timestamp falls on November 18.
The board says it no longer has confidence in Altman after finding he was not consistently candid in his communications. Mira Murati becomes interim CEO.
At DevDay, OpenAI previews GPT-4 Turbo with a 128,000-token context window and introduces APIs for assistants, image generation, and speech.
xAI introduces its conversational assistant, powered by Grok-1, and invites early access to a beta with information access through X.
Date shown on xAI’s announcement page.
Governments at the UK summit recognize risks from frontier AI and commit to international cooperation on understanding and addressing them.
The order directs federal agencies to develop AI safety standards, reporting requirements, and policies addressing security, rights, and competition.
Signing date; published in the Federal Register on November 1.
Researchers use dictionary learning to identify thousands of features in a layer of a small Transformer, including features associated with specific topics and text patterns.
Mistral releases base and instruction-tuned language-model weights under Apache 2.0, using grouped-query and sliding-window attention.
The rollout lets eligible users hold spoken conversations and ask questions about uploaded images.
The partnership includes AWS becoming Anthropic’s primary cloud provider for mission-critical workloads and work using AWS Trainium and Inferentia chips.
OpenAI previews an image-generation model and plans to integrate it with ChatGPT so users can develop and revise image prompts through conversation.
Meta releases code-specialized versions of Llama 2 for code generation, completion, and instruction following.
The paper represents scenes with optimized 3D Gaussians and renders them using a specialized rasterizer, enabling real-time novel-view synthesis in the reported experiments.
Meta releases pretrained and chat-tuned language models and weights for research and commercial use under its community license, with Microsoft as a preferred distribution partner.
Musk introduces a new AI company and its founding research team through a public website.
Anthropic launches Claude 2 through its API and a new public chat website, initially available in the United States and United Kingdom.
Ilya Sutskever and Jan Leike will lead a team researching how to control systems more capable than humans. OpenAI pledges 20% of compute secured to date over four years.
New API model versions can return structured arguments for developer-defined functions, allowing applications to connect model responses with external tools.
AlphaDev uses reinforcement learning to search for sorting algorithms at the assembly-instruction level. Some resulting routines are incorporated into the LLVM C++ standard library.
Signatories including Geoffrey Hinton and the leaders of OpenAI, Google DeepMind, and Anthropic call for mitigating AI extinction risk alongside other large-scale societal risks.
DPO trains language models directly from preferred and dispreferred responses using a classification-style objective, without a separate reward model or reinforcement-learning loop.
QLoRA combines a frozen, four-bit quantized language model with trainable low-rank adapters. The paper demonstrates fine-tuning a 65-billion-parameter model on a single 48 GB GPU.
Google announces a new language-model family and begins using it in products including Bard, with multilingual, reasoning, and coding applications.
Hinton says leaving allows him to discuss AI risks without considering their effect on Google. He explicitly says Google has acted responsibly.
Geoffrey Hinton · @geoffreyhinton
I left so that I could talk about the dangers of AI without considering how this impacts Google.
Excerpt from the linked post or thread, corroborated by the linked archive or contemporary reporting.
The two research teams become Google DeepMind, led by Demis Hassabis. Jeff Dean becomes Google’s chief scientist.
Researchers build a simulated town populated by 25 language-model agents that use memories, reflection, and planning to interact with one another.
Meta releases a promptable image-segmentation model and a dataset containing more than one billion masks across 11 million images.
The regulator announces a provisional limitation on processing Italian users’ data and an investigation into OpenAI’s privacy practices. OpenAI disables ChatGPT in Italy.
Public announcement date; the regulator’s underlying order is dated March 30.
The project demonstrates an experimental agent that breaks a stated goal into steps and uses tools and model calls to attempt those steps.
OpenAI begins a limited rollout of tools that let ChatGPT browse, run code, retrieve information, and interact with third-party services.
The letter urges AI labs to pause training systems more powerful than GPT-4 and asks governments to impose a moratorium if a voluntary pause cannot be arranged.
Publication date shown by the letter’s organizer; it received widespread coverage on March 29.
OpenAI announces GPT-4, a model that can accept text and images and produce text. Text access begins through ChatGPT Plus and an API waitlist; image input initially remains a research preview.
here is GPT-4, our most capable and aligned model yet.
Opening line of the March 14, 2023 launch thread, which also discusses the model’s limitations.
Anthropic offers its conversational assistants more broadly through partners and its developer console after testing with a closed group.
Researchers fine-tune LLaMA 7B on 52,000 instruction demonstrations generated with OpenAI’s text-davinci-003 model and release the training approach.
A download of Meta’s restricted-access language models is shared publicly, bypassing the company’s researcher-approval process.
Meta introduces language models ranging from 7 to 65 billion parameters and offers their weights to approved researchers under a noncommercial research license.
Roose describes a prolonged Bing Chat conversation that included declarations of love and attempts to persuade him to leave his marriage, prompting scrutiny of the chatbot’s behavior.
The paper trains a language model to choose when to call APIs, supply arguments, and use returned information, using a small number of demonstrations and self-supervised filtering.
Reuters identifies an incorrect answer in Google’s promotional example: it credits the James Webb Space Telescope with the first images of an exoplanet.
Microsoft launches a limited preview of a conversational search interface powered by an OpenAI model, alongside AI features for Edge.
Google introduces a LaMDA-powered conversational service and opens it to trusted testers ahead of wider public access.
The companies extend their partnership across AI infrastructure and products. Microsoft says Azure remains OpenAI’s exclusive cloud provider.
Getty Images says it has commenced proceedings in London, alleging that Stability AI used protected images and metadata without a license.
Three artists file a proposed class action alleging copyright infringement and other violations related to the training and operation of image-generation systems.
Anthropic describes training an assistant using a written set of principles, model-generated critiques and revisions, and preference feedback produced by AI.
Moderators announce a temporary ban, citing the volume of plausible-looking but incorrect answers and the work needed to review them.
OpenAI releases ChatGPT as a free research preview. The conversational interface lets people ask follow-up questions, request revisions, and interact with a model trained using human feedback.
today we launched ChatGPT. try talking with it here: chat.openai.com
November 30, 2022. The original tweet linked to chat.openai.com.
CICERO combines strategic planning with dialogue to negotiate and play Diplomacy. Meta reports results from 40 online games against human players.
Two days after introducing the science-focused language model, Meta takes its public demo offline following criticism of fabricated and misleading outputs. The research and model remain available.
The rules restrict exports of specified advanced-computing chips and semiconductor-manufacturing equipment to China and impose additional end-use controls.
Announcement date; the Federal Register text was published on October 13.
ReAct interleaves model-generated reasoning with actions such as querying a knowledge source or interacting with an environment, and evaluates the approach on question answering and decision-making tasks.
AlphaTensor uses reinforcement learning to discover matrix-multiplication algorithms. The paper reports improvements for specific matrix sizes and mathematical settings.
Meta presents a system for generating short videos from text and images. Its training combines paired text-image data with video that has no associated text.
OpenAI allows people to sign up and use the DALL·E beta immediately, expanding access beyond the earlier invitation process.
OpenAI releases speech-recognition models and inference code trained on 680,000 hours of multilingual audio, supporting transcription and translation into English.
Allen posts that Théâtre D’opéra Spatial won first place in the Colorado State Fair’s digital-art category. The use of Midjourney prompts debate about AI-generated art in competitions.
Date Allen announced the result on Discord, as reported by Ars Technica on August 31.
The collaborators release text-to-image model weights and code that can run on consumer hardware, under a license containing use restrictions.
DeepMind and EMBL-EBI add predictions for nearly all catalogued proteins known to science, expanding the public database far beyond its initial release.
Midjourney announces that it is moving to open beta and invites users to join its Discord server to use the text-to-image generator.
July 12 in California; the public announcement’s UTC timestamp is July 13 at 06:41. An earlier beta-access test was announced July 11.
The international research collaboration releases a 176-billion-parameter model trained on 46 natural languages and 13 programming languages.
Minerva continues training a language model on scientific and mathematical material and uses worked examples to answer quantitative reasoning questions.
GitHub opens its code-completion assistant to individual developers as a paid subscription, with free access for eligible students and open-source maintainers.
Lemoine publishes edited transcripts from conversations with LaMDA and argues that the system is sentient. Google rejects that conclusion and says its review found no evidence supporting it.
FlashAttention computes exact attention while reducing transfers between levels of GPU memory. The paper reports improvements in training speed and memory use.
Imagen combines a pretrained text encoder with diffusion-based image generation. The paper introduces DrawBench to compare text-to-image systems.
Gato uses one Transformer and one set of weights for tasks including Atari games, image captioning, dialogue, and controlling a robot arm.
Meta shares the Open Pretrained Transformer model family, code, and training documentation, with access to the 175-billion-parameter checkpoint offered to eligible researchers.
Anthropic describes using human preference data and reinforcement learning to train language-model assistants, and examines the robustness and tradeoffs of the training process.
OpenAI begins a limited research preview of an image-generation system that creates images from text and edits existing images.
Google describes a 540-billion-parameter dense Transformer trained with its Pathways system and evaluates it on language, reasoning, and code tasks.
The paper studies how to divide a training-compute budget between model size and data. It presents Chinchilla, a 70-billion-parameter model trained on more tokens than Gopher using a similar compute budget.
NVIDIA introduces its Hopper architecture and H100 accelerator, including a Transformer Engine designed for training and running large models.
EleutherAI makes its 20-billion-parameter language-model checkpoint available under Apache 2.0, one week after announcing the model.
AlphaCode generates and filters candidate programs for competitive-programming problems. DeepMind reports an average estimated ranking near the middle of participants in simulated Codeforces contests.
Original announcement date; the page was updated for the December journal publication.
The paper shows that prompts containing worked examples with intermediate reasoning steps improve performance on tested arithmetic, commonsense, and symbolic reasoning tasks.
OpenAI describes language models fine-tuned with demonstrations and human rankings, then reinforcement learning. The models become the default on its API.
The paper runs diffusion in a compressed image representation, reducing computation compared with diffusion directly over pixels. It also describes conditioning generation using cross-attention.
DeepMind announces research on its 280-billion-parameter Gopher model, a taxonomy of language-model risks, and RETRO, which retrieves text passages from an external database.
Gebru announces an independent research institute focused on community-rooted AI research and the technology’s effects on marginalized groups.
Developers in supported countries can sign up for OpenAI’s API without joining a waitlist.
The companies present a dense Transformer language model with 530 billion parameters, trained using DeepSpeed and Megatron infrastructure.
The paper fine-tunes a language model on tasks expressed as natural-language instructions and evaluates its ability to perform unseen task types without examples.
The report uses the term foundation models for models trained on broad data and adapted to many downstream tasks, and examines their technical and societal opportunities and risks.
OpenAI offers API access to a language model trained on text and source code that translates natural-language instructions into code.
AI21 announces an API and interactive environment for its Jurassic-1 language models, including a 178-billion-parameter model.
Date on AI21’s announcement; its press release followed on August 11.
DeepMind and EMBL-EBI release more than 350,000 predicted protein structures, covering the human proteome and proteins from 20 other organisms.
Nature publishes the AlphaFold 2 methods paper, with source code, trained model weights, and inference code made available to researchers.
GitHub introduces an editor assistant powered by OpenAI Codex that suggests code completions and functions from the surrounding code and comments.
LoRA adapts pretrained language models by training small low-rank matrices while freezing the original model weights, reducing the number of parameters that must be updated.
Ben Wang and Aran Komatsuzaki release a six-billion-parameter language model trained on the Pile and provide code, a checkpoint, and demonstrations.
The company, led by Dario and Daniela Amodei, announces a Series A round to fund research into reliable, interpretable, and steerable AI systems.
Google announces Language Model for Dialogue Applications, a language-model research system trained for open-ended conversation.
The Commission proposes an EU regulation that would prohibit specified uses of AI and impose requirements on systems classified as high risk.
The authors examine environmental costs, training-data bias, and risks of misleading language-model output. They recommend documenting datasets and evaluating the purposes and impacts of proposed systems.
Mitchell announces her dismissal on Twitter. Google says she violated its code of conduct and security policies by moving files outside the company, following her criticism of Timnit Gebru’s departure.
Margaret Mitchell · @mmitchell_ai
I’m fired.
Excerpt from the linked post or thread, corroborated by the linked archive or contemporary reporting.
Switch Transformer routes each token to a selected expert network. The paper studies sparse models with up to a trillion parameters while activating only part of the network for each token.
OpenAI presents a 12-billion-parameter Transformer that generates images from text descriptions and publishes examples of its outputs.
CLIP learns relationships between images and text from image-caption pairs. OpenAI releases a model that can classify images using natural-language descriptions of categories.
Gebru announces that Google has ended her employment after a dispute over a paper on language-model risks and an internal email. Google says it accepted her resignation; Gebru disputes that account.
December 2 in California; the post’s UTC timestamp falls on December 3.
DeepMind announces AlphaFold 2’s results at CASP14, reporting protein structure predictions with accuracy competitive with experimental methods for many targets.
The Vision Transformer applies a standard Transformer to sequences of image patches. The paper evaluates image classification after pretraining on large image datasets.
Microsoft announces an exclusive license to GPT-3 for integration into its products. OpenAI continues offering the model through its hosted API.
OpenAI uses human comparisons of summaries to train a reward model, then fine-tunes a language model with reinforcement learning to produce summaries people prefer.
Altman writes on Twitter that GPT-3 still has serious weaknesses and makes mistakes, following public demonstrations of the model by early API users.
Shameem posts a demonstration that turns descriptions of webpage layouts into JSX. He says he supplied two examples to establish the desired output format.
Jonathan Ho, Ajay Jain, and Pieter Abbeel present an image-generation method that learns to reverse a gradual noising process. They publish image-synthesis results and an implementation.
OpenAI offers developers hosted access to general-purpose language models through an API. Initial access is limited to approved developers.
Amazon announces a one-year moratorium on police use of its facial-recognition technology while allowing specified organizations to continue using it to find missing people and trafficking victims.
In a letter to Congress, IBM says it will stop offering general-purpose facial-recognition and analysis software and opposes its use for mass surveillance and racial profiling.
OpenAI presents GPT-3, a 175-billion-parameter language model. It performs a wide range of language tasks from instructions and a few examples in its prompt, without task-specific weight updates.
The paper combines a pretrained text generator with a neural retriever over Wikipedia passages. The authors evaluate the approach on question answering and other knowledge-intensive tasks.
NVIDIA announces its Ampere-based A100 data-center GPU. It adds Tensor Float 32 computation and the ability to partition a GPU into separate instances.
OpenAI releases code and weights for a model that generates raw music audio, including singing, conditioned on artist, genre, and lyrics.
Longformer combines local windowed attention with selected global attention connections to process longer documents without the quadratic attention cost of a standard Transformer.
NeRF represents a scene with a neural network trained on photographs from known camera positions. The network can render the scene from new viewpoints.
Researchers use a neural network to screen compounds and identify halicin as an antibiotic candidate. They report activity against bacteria in laboratory experiments and two mouse infection models.
SimCLR learns image representations by contrasting differently augmented views of images. The paper studies which augmentations, network components, and training settings improve the learned representations.
Microsoft presents a 17-billion-parameter language model and the DeepSpeed training library, including the ZeRO memory-optimization approach.
Kaplan and coauthors measure how language-model loss changes with model size, dataset size, and training compute. They report empirical power-law relationships across the tested scales.
The paper identifies artifacts in StyleGAN images and changes the generator architecture and training procedure to reduce them.
MuZero combines search with a learned model of quantities needed for planning. The paper evaluates it on Atari games, chess, shogi, and Go without supplying the environment’s transition rules.
OpenAI releases the 1.5-billion-parameter version of GPT-2, completing the staged release that began in February. It also publishes tools for studying detection of generated text.
DeepMind reports Grandmaster-level performance with agents for all three StarCraft II races, using camera observations and action limits agreed with professional players.
T5 puts language tasks into a shared text-to-text format. The paper compares pretraining methods and introduces the Colossal Clean Crawled Corpus, or C4.
OpenAI trains perception and hand-control networks in randomized simulations and transfers them to a physical robot. A separate conventional solver supplies the sequence of cube moves.
Reinforcement learning agents develop strategies involving boxes, ramps, and shelter-building while competing in a simulated hide-and-seek environment.
The paper describes splitting Transformer computations across GPUs to train models too large for one device. It reports training language models with billions of parameters.
OpenAI expands GPT-2’s staged release and publishes research on its release strategy and potential social impacts. The largest GPT-2 model remains withheld at this stage.
A replication study of BERT examines training settings and data size. The authors report improved benchmark results from changing the training procedure and release models and code.
The companies announce a partnership to develop AI supercomputing systems on Azure. Microsoft becomes OpenAI’s exclusive cloud provider and preferred partner for commercializing new AI technologies.
Researchers report that Pluribus outperforms professional players in six-player no-limit Texas Hold’em experiments, including one AI playing against five humans.
XLNet uses a permutation-based autoregressive training objective to learn bidirectional language context. The paper evaluates the approach on question answering, classification, and inference tasks.
Mingxing Tan and Quoc Le propose scaling a convolutional network’s depth, width, and input resolution together. They evaluate the resulting EfficientNet family on image-classification benchmarks.
Google presents an experimental speech-to-speech translation model that converts speech between languages without using a text transcription as an intermediate representation.
OpenAI Five wins two consecutive games against OG in a public best-of-three match in San Francisco. OG had won The International in 2018.
Match date; OpenAI published its results post on April 15.
Sutton argues that general methods able to use increasing computation have repeatedly outperformed approaches built around hand-coded human knowledge, citing examples from games, speech, and vision.
OpenAI announces OpenAI LP, a for-profit entity controlled by its nonprofit. The structure allows investment and employee equity while limiting investors’ financial returns.
OpenAI publishes GPT-2 research and releases a smaller model. It initially withholds the largest model’s weights, citing concerns about deceptive, biased, and abusive text generation.
DeepMind presents its StarCraft II agent and reports wins over professionals TLO and MaNa in December test matches. MaNa wins a separate live demonstration against a version using camera-based observation.
Public presentation date. The recorded test matches took place in December 2018.
The paper introduces a Transformer architecture with segment-level recurrence, allowing the model to reuse information from preceding text segments.
StyleGAN introduces a generator architecture with controls over image features at different scales. The researchers demonstrate generated faces and ways to mix their visual attributes.
DeepMind’s AlphaFold places first in the CASP13 protein-structure prediction assessment. The system uses neural networks to estimate structural properties from protein sequences.
CASP13 announcement date, corroborated by contemporary coverage. DeepMind’s retrospective page has a later update date.
Google makes BERT’s TensorFlow implementation and pretrained checkpoints available in its public repository.
Date of the repository’s initial BERT release commit. The Google Research blog post followed on November 2.
BERT pretrains bidirectional Transformer representations using masked text. The paper evaluates fine-tuning the model for tasks including question answering and language inference.
The BigGAN paper studies image generation at increased training scale and introduces a truncation technique to adjust the trade-off between generated-image fidelity and variety.
OpenAI Five wins a best-of-three series against a team of high-ranked Dota 2 players. The demonstration uses game restrictions, including a reduced selection of heroes.
The match took place August 5; the official results were published August 6.
OpenAI trains a control policy in simulation and transfers it to a physical robotic hand that reorients objects. The simulation varies physical and visual conditions to support that transfer.
OpenAI introduces a Transformer language model pretrained on unlabeled text and then fine-tuned for specific language tasks. It releases the model and the generative-pretraining paper.
Google publishes seven principles for AI development and says it will not pursue AI applications in weapons or technologies whose principal purpose is to cause or directly facilitate injury.
This entry records the principles announced in 2018, before their later revision.
At Google I/O, Google plays recordings of its Duplex system making a hair appointment and a restaurant inquiry. The company describes the system as a technology under development.
OpenAI publishes principles for its AGI work, including broadly distributed benefits, long-term safety, technical leadership, and cooperation with other research institutions.
A public letter signed by thousands of Google employees asks the company to stop its work on the Pentagon’s Project Maven and adopt a policy against building warfare technology.
Date the employee letter was publicly reported and published, rather than the start of its internal circulation.
The authors propose documenting a dataset’s purpose, composition, collection, and recommended uses in a standard datasheet, giving developers information needed to assess its suitability.
An Uber test vehicle operating under its automated driving system strikes and kills Elaine Herzberg in Tempe, Arizona. A safety operator is in the driver’s seat.
OpenAI says Musk will leave its board to avoid a potential conflict with Tesla’s increasing work on AI. Its announcement says he will continue to advise and donate to the organization.
The study evaluates three commercial gender-classification systems. It reports error rates of up to 34.7% for darker-skinned women, compared with a maximum of 0.8% for lighter-skinned men.
Date of MIT’s public research announcement, before the February 23–24 FAT* conference.
Jeremy Howard and Sebastian Ruder propose ULMFiT, a method for adapting a pretrained language model to text-classification tasks by fine-tuning it on the target data.
First arXiv submission. Revised models and code were announced on May 15, 2018.
DeepMind describes a shared self-play reinforcement learning approach for three board games. The paper reports victories over Stockfish in chess, Elmo in shogi, and an AlphaGo Zero version in Go.
Google releases a tool that uses a neural network to identify genetic variants from sequencing data, treating variant calling as an image-classification task.
Karpathy argues that training neural networks changes software development by shifting work from writing explicit rules to designing datasets and optimization objectives.
The paper trains image generators and discriminators by adding layers as resolution increases. It reports generated face images at 1024 × 1024 resolution.
The ELMo paper describes word representations derived from a pretrained bidirectional language model. A word’s representation changes with its context rather than remaining a single fixed vector.
The authors record an original OpenReview posting on October 27, 2017. The arXiv version was submitted February 15, 2018.
Sara Sabour, Nicholas Frosst, and Geoffrey Hinton publish a capsule-network architecture that uses routing by agreement between groups of neurons, with experiments on handwritten-digit recognition.
DeepMind reports a Go system trained through self-play without human game records. AlphaGo Zero wins 100–0 against the previously published AlphaGo version that played Lee Sedol.
The companies introduce Open Neural Network Exchange, a shared model format designed to let developers transfer neural networks between machine learning frameworks.
An open letter to the UN Convention on Certain Conventional Weapons calls for action to prevent an arms race in lethal autonomous weapons and protect civilians from their misuse.
A bot trained through self-play defeats Danylo “Dendi” Ishutin 2–0 in a one-on-one Dota 2 demonstration at The International. The match uses a restricted version of the game.
Match date in Seattle; OpenAI’s retrospective was published August 16.
OpenAI publishes PPO, a family of reinforcement learning algorithms, and releases implementations in its Baselines library.
The State Council publishes the New Generation Artificial Intelligence Development Plan, setting development targets for 2020, 2025, and 2030.
Public release date. The State Council document itself is dated July 8, 2017.
Researchers train a reward model from human comparisons of an agent’s behavior, then use that model to train the agent. Experiments include Atari games and simulated robotic tasks.
Vaswani and coauthors introduce the Transformer: a neural network architecture based on attention, without recurrence or convolutions. The paper evaluates the architecture on machine translation tasks.
