The difficult part was not getting every cell filled. It was deciding when the available evidence supported a classification, and when it did not.
The records came from public announcements, so the underlying information was uneven. Some entries carried detailed financing information; others had a name and very little else.
The issues I kept working through were missing values, entity types that were inconsistent or unclear, categories that applied to some organizations but not others, incomplete financing information, ambiguous stage evidence, and records that should be marked unknown or not applicable rather than pushed into a category.
Different working versions of the dataset also contained different record counts and units of analysis. Rather than treating those counts as directly comparable, I focused on standardizing the records needed for the next analytical stage.