If we have a real physical object of height 1, distance 1 from the lens. and let the projected height be h on the focal plane.
Then, we see that 1/1=h/f, so h=f. In general, we will have that h∝f, so the focal length is proportional to the zoom factor.
So, 'Zoom lenses' (Which support multiple focal lengths) report the zoom factor in terms of focal length. For example, a 24-70mm lens has a zoom factor of 70/24 = 2.9x.
§ Focal Length Photographically (2. Field of View)
Suppose we have a fixed screen height w. Now, we want to understand what angle of vision we can capture on the screen height w.
We have tan(theta/2)=(w/2)/f. Thus, theta=2arctan(w/(2f)). So, theta decreases as f increases.
So, if we want a wide angle lens, it must have a small focal length.
Easy way to see this: keep a triangle with a fixed base (this is the sensor) at y=0. Move the apex of the triangle (this is the lens) closer to the base. As we do this, the angle at the apex increases.
The angle at the apex, when extended, corresponds to the 'extent of the object' that can be captured on the sensor.
Two extreme cases: When the apex 'lies on' the base, the angle is 180 degrees, and the object can be as large as we want.
Another extreme case: When the apex is very very far away, the angle is very small, and the object can be very small.
The aperture is measured in units of f/k, where f is the focal length.
The units are called 'f-stops'.
Why stops? Well, cause a dude called 'Waterhouse' designed an interchangeable 'stop' (a part of an optics device that can stop light) to control the amount of light that enters the camera.
Now, while aperture itself (physically) is measured in diameters, what it lets in (light) has units of area. So, increasing the diameter by 2 doubles the light.
Thus, f-stops are available in f/1.4, f/2, f/2.8, where each has an aperture that is 2 smaller than the previous one. Consequently, they let in half the light of the previous one.
§ Aperture Photographically (Depth of Field, Maximum Circle of Confusion)
My intuition is the following. If we have a pinhole camera, then everything is sharp, since we get a 'single' point of light from each point in the scene. But this corresponds to aperture of ϵ (very small).
As we increase the aperture, we get a cone of light from each point in the scene.
This means that not everything lies on the focal plane.