Yeah i agree. When the child is shown the labeled example of an elephant and infers the traits that make an elephant an elephant, her previous visual experiences provide background knowledge that restricts the space of hypotheses she considers. After all, there's an infinite set of logically consistent hypotheses.
Nevertheless, if you provide your machine system with the video of all the child's visual input, it still won't generalize well from single examples, the way children do effortlessly.
Humans are excellent at learning from very few or even 1 example. Show a toddler a single image of an elephant and the toddler will generalize perfectly on new examples; show a machine a few thousand images of elephants and it might generalize decently if your machine is really clever.
There are very few tasks where machine systems achieve anything resembling human level performance. But on all such tasks, the machine requires far more data and still underperforms.
Nevertheless, if you provide your machine system with the video of all the child's visual input, it still won't generalize well from single examples, the way children do effortlessly.