Structure discovery in conditional probability distributions via an entropic prior and parameter extinction

    •  Brand, M.E., "Structure Discovery in Conditional Probability Distributions via an Entropic Prior and Parameter Extinction", Neural Computation Journal, Vol. 11, No. 5, pp. 1155-1182, July 1999.
      BibTeX Download PDF
      • @article{Brand1999jul,
      • author = {Brand, M.E.},
      • title = {Structure Discovery in Conditional Probability Distributions via an Entropic Prior and Parameter Extinction},
      • journal = {Neural Computation Journal},
      • year = 1999,
      • volume = 11,
      • number = 5,
      • pages = {1155--1182},
      • month = jul,
      • url = {}
      • }
  • MERL Contact:

We introduce an entropic prior for multinomial parameter estimation problems and solve for its maximum a posteriori (MAP) estimator. The prior is a bias for maximally structured and minimally ambiguous models. In conditional probability models with hidden state, iterative MAP estimation drives weakly supported parameters toward extinction, effectively turning them off. Thus structure discovery is folded into parameter estimation. We then establish criteria for simplifying a probabilistic model's graphical structure by trimming parameters and states, with a guarantee that any such deletion will increase the posterior probability of the model. Trimming accelerates learning by sparsifying the model. All operations monotonically and maximally increase the posterior probability, yielding structure-learning algorithms only slightly slower than parameter estimation via expectation-maximization (EM), and orders of magnitude faster than search-based structure induction. When applied to hidden Markov model (HMM) training, the resulting models show superior generalization to held-out test data. In many cases the resulting models are so sparse and concise that they are interpretable, with hidden states that strongly correlate with meaningful categories.