We propose a method to infer skeleton animation by using a single image. CNN is used to learn the image of the model. Using the learned data, a total of 13 joint angles including the root are deduced to generate a skeleton animation. We have experimented with black and white images made using a pre-prepared 3D model. By using the single channel of black and white image, learning can be done more easily and quickly. A single image is processed through preprocessing. Through learning, the joints of the character are extracted and remapped to the model using the extracted joints to check how similar they are. We show that it is possible and possible to capture human motion only through learning of this single image. Motion capture using machine learning is expected to solve the problems of high cost, low efficiency, current motion capture and animation production. In addition, it is expected that it will be extended to various derivative studies rather than learning using the relationship between simple images and joints.