All Skills
25,836 skillsDemystifying Clip Data Arxiv 2309 16671v4
Demystifying CLIP Data
6
Gpt 4vision System Card Openai Gpt4v 2023
GPT-4V(ision) System Card
6
Visual Instruction Tuning Arxiv 2304 08485v2
Visual Instruction Tuning
6
Instruction Tuning With Gpt 4 Arxiv 2304 03277v1
Instruction Tuning with GPT-4
6
Lima Less Is More For Alignment Arxiv 2305 11206v1
LIMA: Less Is More for Alignment
6
Curriculum Learning Crossref Icml 2009 Curriculum
Curriculum Learning
6
Phi 1 Textbooks Are All You Need Arxiv 2306 11644v2
Phi-1: Textbooks Are All You Need
6
Llama 3 The Llama 3 Herd Of Models Arxiv 2407 21783v2
Llama 3: The Llama 3 Herd of Models
6
Matryoshka Representation Learning Arxiv 2205 13147v4
Matryoshka Representation Learning
6
Snli Ve Visual Entailment Dataset Arxiv 1901 06706v1
SNLI-VE: Visual Entailment Dataset
6
Influence Functions In Deep Learning Arxiv 2002 08484v3
Influence Functions in Deep Learning
6
Phi 15 Textbooks Are All You Need Ii Arxiv 2309 05463v2
Phi-1.5: Textbooks Are All You Need II
6
Imagenet 21k Pretraining For The Masses Arxiv 2104 10972v4
ImageNet-21K Pretraining for the Masses
6
Pixtral 12b A Frontier Multimodal Model Arxiv Pixtral 2024
Pixtral 12B: A Frontier Multimodal Model
6
Scaling Laws For Neural Language Models Arxiv 2001 08361v1
Scaling Laws for Neural Language Models
6
Deita What Makes Good Data For Alignment Arxiv 2312 15685v2
Deita: What Makes Good Data for Alignment
6
Scaling Data Constrained Language Models Arxiv 2305 16264v3
Scaling Data-Constrained Language Models
6
Trak Attributing Model Behavior At Scale Arxiv 2303 14186v2
TRAK: Attributing Model Behavior at Scale
6
Nocaps Novel Object Captioning At Scale Arxiv 1812 08658v2
Nocaps: Novel Object Captioning at Scale
6
Idefics2 An 8b Parameters Multimodal Model Arxiv 2405 02246v
Idefics2: An 8B Parameters Multimodal Model
6
Dall E Zero Shot Text To Image Generation Arxiv 2102 12092v2
DALL-E: Zero-Shot Text-to-Image Generation
6
Grit General Robust Image Task Benchmark Arxiv 2306 14818v2
Grit: General Robust Image Task Benchmark
6
Minicpm V A Gpt 4v Level Mllm On Your Phone Arxiv 2408 01800
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
6
Sigmoid Loss For Language Image Pre Training Arxiv 2303 1534
Sigmoid Loss for Language Image Pre-Training
6
Data Centric Artificial Intelligence A Survey Arxiv 2303 101
Data-Centric Artificial Intelligence: A Survey
6
Scaling Instruction Finetuned Language Models Arxiv 2210 114
Scaling Instruction-Finetuned Language Models
6
Scaling Vision With Sparse Mixture Of Experts Arxiv 2106 059
Scaling Vision with Sparse Mixture of Experts
6
Emu Generative Pretraining In Multimodality Arxiv 2307 05222
Emu: Generative Pretraining in Multimodality
6
Sa 1b Segment Anything 1 Billion Masks Dataset Arxiv Sa1b 20
SA-1B: Segment Anything 1 Billion Masks Dataset
6
Webvid 10m A Large Scale Video Text Dataset Arxiv 2104 00650
WebVid-10M: A Large-Scale Video-Text Dataset
6
Hard Negative Mixing For Contrastive Learning Arxiv 2010 010
Hard Negative Mixing for Contrastive Learning
6
Coco Microsoft Coco Common Objects In Context Arxiv 1405 031
COCO: Microsoft COCO: Common Objects in Context
6
Open Vocabulary Object Detection Using Captions Arxiv 2011 1
Open-Vocabulary Object Detection Using Captions
6
Capsfusion Rethinking Image Text Data At Scale Arxiv 2310 20
CapsFusion: Rethinking Image-Text Data at Scale
6
Dolphins Multimodal Language Model For Driving Arxiv 2312 00
Dolphins: Multimodal Language Model for Driving
6
Masked Autoencoders Are Scalable Vision Learners Arxiv 2111
Masked Autoencoders Are Scalable Vision Learners
6