Training Set Data
Given the training set, the next task is to generate data for developing a force field. Quantum mechanics calculations are powerful for producing large amounts of information in terms of energy, energy derivatives, electrostatic potentials and electron densities, atomic partial charges, and so on, which can be used to derive force field parameters. However, how to sample the molecular conformations is a critical issue. Only with sufficient sampling of the conformational spaces can the force field derived be expected to represent the conformers well.
Using small fragments implies that long-range (nonbond) interactions are to be considered separately. In addition, the nonbond interactions cannot be represented well in QM calculations. Even when electron correlation is considered, polarization and multibody effects are very difficult to capture using QM calculations. A common practice is to use experimental data of molecular liquids to optimize the nonbond interaction parameters. Therefore, another training set of liquid molecules, which often overlaps with the training set of fragments, is required. The experimental data usually includes equilibrium data such as liquid density at given temperature and pressure, heat of vaporization, and vapor–liquid equilibrium curves, as well as dynamic data such as diffusion coefficients and viscosity.