TY - GEN
T1 - Declarative probabilistic programming with datalog
AU - Barany, Vince
AU - Cate, Balder Ten
AU - Kimelfeld, Benny
AU - Olteanu, Dan
AU - Vagena, Zografoula
N1 - Funding Information: We are thankful to Molham Aref, Todd J. Green and Emir Pasalic for insightful discussions and feedback on this work. We also thank Michael Benedikt, Georg Gottlob and Yannis Kassios for providing useful comments and suggestions. Benny Kimelfeld is a Taub Fellow, supported by the Taub Foundation. Kimelfeld's work was supported in part by the Israel Science Foundation. Dan Olteanu acknowledges the support of the EPSRC programme grant VADA. This work was supported by DARPA under agreements #FA8750-15-2-0009 and #FA8750-14-2-0217. The U.S. Government is authorized to reproduce and distribute reprints for governmental purposes notwithstanding any copyright thereon. Publisher Copyright: © 2016 Vince Barany, Balder ten Cate, Benny Kimelfeld, Dan Olteanu, and Zografoula Vagena.
PY - 2016/3/1
Y1 - 2016/3/1
N2 - Probabilistic programming languages are used for developing statistical models, and they typically consist of two components: a specification of a stochastic process (the prior), and a specification of observations that restrict the probability space to a conditional subspace (the posterior). Use cases of such formalisms include the development of algorithms in machine learning and artificial intelligence. We propose and investigate an extension of Datalog for specifying statistical models, and establish a declarative probabilistic-programming paradigm over databases. Our proposed extension provides convenient mechanisms to include common numerical probability functions; in particular, conclusions of rules may contain values drawn from such functions. The semantics of a program is a probability distribution over the possible outcomes of the input database with respect to the program. Observations are naturally incorporated by means of integrity constraints over the extensional and intensional relations. The resulting semantics is robust under different chases and invariant to rewritings that preserve logical equivalence.
AB - Probabilistic programming languages are used for developing statistical models, and they typically consist of two components: a specification of a stochastic process (the prior), and a specification of observations that restrict the probability space to a conditional subspace (the posterior). Use cases of such formalisms include the development of algorithms in machine learning and artificial intelligence. We propose and investigate an extension of Datalog for specifying statistical models, and establish a declarative probabilistic-programming paradigm over databases. Our proposed extension provides convenient mechanisms to include common numerical probability functions; in particular, conclusions of rules may contain values drawn from such functions. The semantics of a program is a probability distribution over the possible outcomes of the input database with respect to the program. Observations are naturally incorporated by means of integrity constraints over the extensional and intensional relations. The resulting semantics is robust under different chases and invariant to rewritings that preserve logical equivalence.
KW - Chase
KW - Datalog
KW - Probabilistic Programming
KW - Probability Measure Space
UR - https://www.scopus.com/pages/publications/84962741006
U2 - 10.4230/LIPIcs.ICDT.2016.7
DO - 10.4230/LIPIcs.ICDT.2016.7
M3 - Conference contribution
T3 - Leibniz International Proceedings in Informatics, LIPIcs
BT - 19th International Conference on Database Theory, ICDT 2016
A2 - Martens, Wim
A2 - Zeume, Thomas
T2 - 19th International Conference on Database Theory, ICDT 2016
Y2 - 15 March 2016 through 18 March 2016
ER -