{"id":1,"slug":"ilya-sutskever","name":"Ilya Sutskever","title":"Co-Founder & Chief Scientist","company":"Safe Superintelligence Inc.","sector":"general","profile_url":"https://en.wikipedia.org/wiki/Ilya_Sutskever","image_url":"/api/v1/ceo-ai-leaderboard/portrait/ilya-sutskever.jpg","score":97,"tier":"frontier_builder","dimensions":{"foundations":20,"vector_embeddings":19,"transformers_lm":20,"frontier_founder":20,"lm_domain_depth":19,"hands_on_engineering":20,"industry_impact":20,"scientific_founder":17},"rubric_version":3,"weighted_score":97,"penalties":{"bought_popularity":0,"capital_without_competence":0},"rationale":"Sutskever earned a PhD in computer science at the University of Toronto (thesis: 'Training Recurrent Neural Networks', 2013) under Geoffrey Hinton, and personally co-authored AlexNet (2012, with Krizhevsky and Hinton) which catalyzed the deep-learning era. He co-invented sequence-to-sequence learning with attention-adjacent architectures (Sutskever, Vinyals, Le 2014), a direct precursor in the seq2seq->transformer lineage, and was a co-author on 'Distributed Representations of Words and Phrases' (word2vec, 2013). As OpenAI co-founder and chief scientist (2015-2024) he personally shaped GPT-2/GPT-3/GPT-4 research direction and post-training. This is a canonical, field-defining research and engineering record spanning math foundations through the full attention/transformer/scaling lineage, not organizational leadership alone.\n\nSutskever authored building blocks that today's frontier models directly descend from: 'Sequence to Sequence Learning with Neural Networks' (2014) is the encoder-decoder precursor the transformer displaced yet built on, 'Distributed Representations of Words and Phrases' (word2vec, 2013) is canonical embedding work, and as OpenAI chief scientist he co-authored 'Language Models are Few-Shot Learners' (GPT-3, 2020) and CLIP (2021) — all cited by and built into GPT/Claude/Gemini/Llama-class systems. His language-modeling record is continuous from pre-word2vec neural LMs ('Generating Text with Recurrent Neural Networks', ICML 2011) and his RNN-training PhD thesis (2013) through seq2seq, GPT-2/3/4 pretraining and alignment, and now SSI — ~15 years, still active. As a technical founder he co-founded OpenAI (2015) serving as chief scientist personally setting research direction through May 2024 (~9 years), then co-founded and now leads Safe Superintelligence (2024–), ~11 years total across two such companies.","evidence":[{"claim":"PhD in computer science, University of Toronto, 2013, advisor Geoffrey Hinton, thesis 'Training Recurrent Neural Networks'","source_url":"https://en.wikipedia.org/wiki/Ilya_Sutskever","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"Co-inventor of AlexNet with Alex Krizhevsky and Geoffrey Hinton (2012 ImageNet paper, 200k+ citations on Google Scholar)","source_url":"https://scholar.google.com/citations?user=x04W_mMAAAAJ&hl=en","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"Co-author 'Distributed Representations of Words and Phrases and their Compositionality' (word2vec extension, 2013)","source_url":"https://doi.org/10.48550/arxiv.1310.4546","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"Co-author 'Sequence to Sequence Learning with Neural Networks' (2014), a foundational seq2seq paper in the pre-transformer attention lineage","source_url":"https://doi.org/10.48550/arxiv.1409.3215","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"OpenAI co-founder (2015) and Chief Scientist through May 2024, overseeing GPT research; now CEO/co-founder of Safe Superintelligence Inc.","source_url":"https://www.cnbc.com/2025/07/03/ilya-sutskever-is-ceo-of-safe-superintelligence-after-meta-hired-gross.html","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"PhD under Geoffrey Hinton at University of Toronto; specializes in machine learning; co-created AlexNet with Krizhevsky and Hinton; won NeurIPS Test of Time Award three years running (2022-2024)","source_url":"https://en.wikipedia.org/wiki/Ilya_Sutskever","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"Google Scholar profile (x04W_mMAAAAJ) lists ~848,637 citations, h-index 109, i10-index 172; top works ImageNet/AlexNet (2012), Language Models are Few-Shot Learners (2020), CLIP (2021), Dropout (2014), Sequence to Sequence Learning with Neural Networks (2014)","source_url":"https://scholar.google.com/citations?user=x04W_mMAAAAJ&hl=en","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"Sutskever, Vinyals and Le won the NeurIPS 2024 Test of Time award for 'Sequence to Sequence Learning with Neural Networks'","source_url":"https://blog.neurips.cc/2024/11/27/announcing-the-neurips-2024-test-of-time-paper-awards/","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"Test of Time award talk for 'Distributed Representations of Words and Phrases and their Compositionality' (word2vec) at NeurIPS 2023","source_url":"https://neurips.cc/virtual/2023/test-of-time/83333","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"Theory/foundations papers authored with Hinton: 'Deep, narrow sigmoid belief networks are universal approximators' (Neural Comput, 2008) and 'Temporal-kernel recurrent neural networks' (Neural Netw, 2010)","source_url":"https://pubmed.ncbi.nlm.nih.gov/18533819/","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"Sequence to Sequence Learning with Neural Networks (2014) — foundational encoder-decoder work in the seq2seq→transformer lineage; NeurIPS 2024 Test of Time award","source_url":"https://doi.org/10.48550/arxiv.1409.3215","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"Language Models are Few-Shot Learners (GPT-3, 2020) — a frontier-model paper Sutskever co-authored as OpenAI chief scientist","source_url":"https://doi.org/10.48550/arxiv.2005.14165","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"Generating Text with Recurrent Neural Networks (ICML 2011) — pre-word2vec neural language-modeling work, anchoring 15 years of continuous LM research","source_url":"https://icml.cc/2011/papers/524_icmlpaper.pdf","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"OpenAI co-founder (2015) and Chief Scientist through May 2024; co-founder and CEO of Safe Superintelligence Inc. (2024–)","source_url":"https://en.wikipedia.org/wiki/Ilya_Sutskever","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"Sequence to Sequence Learning with Neural Networks (Sutskever, Vinyals, Le, 2014) — direct precursor in the seq2seq→transformer lineage frontier LMs descend from; NeurIPS 2024 Test of Time","source_url":"https://doi.org/10.48550/arxiv.1409.3215","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"Co-author of GPT-3 'Language Models are Few-Shot Learners' (2020) and word2vec 'Distributed Representations of Words and Phrases' (2013) — both directly built into the frontier LM stack","source_url":"https://doi.org/10.48550/arxiv.1310.4546","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"},{"claim":"OpenAI co-founder (2016 per Wikidata) and Chief Scientist through 2024; now co-founder of Safe Superintelligence Inc. — technical/scientific founder authoring core research","source_url":"https://en.wikipedia.org/wiki/Ilya_Sutskever","verified":true,"verified_at":"2026-09-14T02:18:59.774252+00:00"}],"confidence":0.96,"source":"seeded","status":"published","scored_at":"2026-09-14T02:18:59.774252+00:00","rank":1,"sector_rank":1,"penalty_evidence":[],"metadata":{"education":["PhD Computer Science, University of Toronto (2013, advisor Geoffrey Hinton)","BSc/MSc, University of Toronto/Open University of Israel"],"canonical_papers":["ImageNet Classification with Deep Convolutional Neural Networks (AlexNet, 2012)","Sequence to Sequence Learning with Neural Networks (2014)","Distributed Representations of Words and Phrases and their Compositionality (2013)","Dropout: A Simple Way to Prevent Neural Networks from Overfitting (2014)"],"first_verifiable_year":2007,"notable_systems":["AlexNet","OpenAI GPT-2/GPT-3/GPT-4 research direction","AlphaGo (co-author on Nature paper)","Safe Superintelligence Inc."],"citations":219277,"h_index":62,"patents":0,"dossier_notes":"Dossier's OpenAlex figures (h-index 62, 219k citations) are conservative relative to the live Google Scholar profile (h-index 109, 848k+ citations) — OpenAlex undercounts; both sources agree on canonical works. No homonym risk; PubMed sample entries (Hinton co-authorship) match the correct person.","years_language_modeling":15,"years_as_technical_founder":11,"frontier_lineage":["seq2seq encoder-decoder architecture (2014)","word2vec distributed word embeddings (2013)","GPT-3 few-shot language modeling (2020)","CLIP contrastive vision-language pretraining (2021)","GPT-2/3/4 pretraining and alignment research direction"],"technical_founder_roles":["OpenAI — co-founder & Chief Scientist — 2015–2024 (~9 yrs)","Safe Superintelligence Inc. — co-founder & CEO — 2024–2026 (~2 yrs)"]},"dimension_labels":{"foundations":"Mathematical Foundations","vector_embeddings":"Vector Embeddings","transformers_lm":"Transformer & LM Lineage","frontier_founder":"Frontier Founder","lm_domain_depth":"Deep Knowledge Domain Expert","hands_on_engineering":"Hands-On Engineering","industry_impact":"Scientific & Industry Impact","scientific_founder":"Scientific & Technical Founder"},"passes":[{"pass":"pass_1","dimensions":{"frontier_founder":20,"lm_domain_depth":19,"scientific_founder":17},"confidence":0.93,"duration_ms":44593},{"pass":"pass_2","dimensions":{"frontier_founder":20,"lm_domain_depth":19,"scientific_founder":17},"confidence":0.93,"duration_ms":41484}],"dossier_sources":{},"validation":{"recompute":"weighted_score = round(70 * (foundations + vector_embeddings + transformers_lm + frontier_founder + lm_domain_depth) / 100 + 30 * (hands_on_engineering + industry_impact + scientific_founder) / 60)","score_formula":"score = max(0, weighted_score - bought_popularity - capital_without_competence)","methodology":"/api/v1/ceo-ai-leaderboard/methodology","export":"/api/v1/ceo-ai-leaderboard/export.json"}}