Saurav Kadavath

According to our database¹, Saurav Kadavath authored at least 15 papers between 2019 and 2023.

Collaborative distances:

Dijkstra number² of four.
Erdős number³ of three.

Timeline

Legend:

Book

In proceedings

Article

PhD thesis

Dataset

Other

Links

On csauthors.net:

Bibliography

2023

Specific versus General Principles for Constitutional AI.

[BibT_eX]

[DOI]

CoRR, 2023

Measuring Faithfulness in Chain-of-Thought Reasoning.

[BibT_eX]

[DOI]

CoRR, 2023

The Capacity for Moral Self-Correction in Large Language Models.

[BibT_eX]

[DOI]

CoRR, 2023

Discovering Language Model Behaviors with Model-Written Evaluations.

[BibT_eX]

[DOI]

Timothy Telleen-Lawton

Proceedings of the Findings of the Association for Computational Linguistics: ACL 2023, 2023

2022

Discovering Language Model Behaviors with Model-Written Evaluations.

[BibT_eX]

[DOI]

Timothy Telleen-Lawton

CoRR, 2022

Constitutional AI: Harmlessness from AI Feedback.

[BibT_eX]

[DOI]

CoRR, 2022

DeepChrome 2.0: Investigating and Improving Architectures, Visualizations, & Experiments.

[BibT_eX]

[DOI]

Saurav Kadavath

Samuel Paradis

Jacob Yeung

CoRR, 2022

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

[BibT_eX]

[DOI]

CoRR, 2022

Language Models (Mostly) Know What They Know.

[BibT_eX]

[DOI]

CoRR, 2022

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

[BibT_eX]

[DOI]

CoRR, 2022

2021

Pretraining & Reinforcement Learning: Sharpening the Axe Before Cutting the Tree.

[BibT_eX]

[DOI]

Saurav Kadavath

Samuel Paradis

Brian Yao

CoRR, 2021

Measuring Coding Challenge Competence With APPS.

[BibT_eX]

[DOI]

Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, 2021

Measuring Mathematical Problem Solving With the MATH Dataset.

[BibT_eX]

[DOI]

Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, 2021

The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization.

[BibT_eX]

[DOI]

Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision, 2021

2019

Using Self-Supervised Learning Can Improve Model Robustness and Uncertainty.

[BibT_eX]

[DOI]

Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, 2019

Saurav Kadavath

Timeline

Legend:

Links

On csauthors.net:

Bibliography

Loading...