Dan Heller's Photography Business Blog Industry analysis from www.danheller.com

The photography world -- the business, the culture, the art, the politics, the technology.

Site Feed

Subscribe to
Posts [Atom]

View mobile version

My Photo
Name:
Location: Santa Cruz, California, United States
My Books on the
Photography Business

Thursday, March 15, 2012

Pinterest Copyright Infringement: Yeah, so what?

The latest hot startup in the photo-sharing space is one that is also creating a lot of controversy about copyright infringement. Pinterest lets users create "boards" of images they find from around the Web. Users “pin up” these images, and share them with friends and strangers.


“Is this copyright infringement,” you ask?


Well, imagine exactly the same website that let's users upload music or movies. Do you think the music labels or movie studios would permit this? Pinterest would be shut down before they could get their first dollar of venture capital.


“But they’re photos, not music or movies!”


Yes, and photos have precisely the same copyright protection.


“Ok, wise guy, then why hasn’t Pinterest been shut down?”


Simply put, there’s no one there to stop them, at least not with the same effect and scale as music labels or movie studios. And the reason is reflected in your very statement: society simply regards photos as “different” from movies; they don’t see Pinterest’s use of imagery as copyright infringement.


And this is a natural feeling to the common person. Everyone shoots pictures all the time; it takes milliseconds; most people don’t invest any thought or intent. By contrast, music and movies require considerable time, effort and expense to produce. So, there’s a difference.


And herein lies the unresolved problem: the law is the law, and photos are copyrighted works, regardless of the time, skill, or anything else necessary to create them. Accordingly, photos are supposed to enjoy the same legal protections as music and movies.


“I see. But, most people want and expect to share their images with others.”


Yes! Their images. Pinterest isn’t letting people share their own photos; they’re sharing other people’s photos.”


“Ah, I see now.”


Very good, Grasshopper.


As a society, we permit this kind of infringement, which explains why there are no large, powerful, influential organizations representing the interests (and the copyrights) of photographers. People simply regard photos as different.


A case in point can be found in this article on chow.com, discussing people’s reactions when they found their recipes were being “pinned” to Pinterest, along with the photos of their foods. The complaint was that their intellectual property (cooking recipes) were being stolen; the recommendation: “Just allow the photo to be shared, not my recipe!”


You see? Never mind the pro photographers whose pictures were being infringed; they’re not part of the conversation.


“Ok, so what about those professional photographers? How are they hurt?”


I’ve been a photo industry analyst since the mid 1990s, and I’ve seen the industry suffer more from “piracy” than the film and music industries combined. Every single publicly traded stock photo agency has either gone out of business or withdrawn from public trade. Getty Images is the last profitable company of any significant size, and even then, its pay to photographers has been drifting lower for over ten years to maintain that status. A series of studies from Picscout – a photo-tracking service for stock agencies and photographers – finds that 90% of commercial websites use at least one photo in a manner considered to be “commercial use” without the copyright holder's authorization. No company whose business model is to sell or license photography has had venture capital investment since 2000.


Yet, the shadow economy for photography is enormous. In a study I conducted in 2007 on contract for a potential investor in a photo-related technology, I found that most photo buying and licensing was done on a peer-to-peer basis, mostly in local markets and exchanges, at a scale that suggested the total economic activity tipped at $25B/year. Yet, none of it can migrate online because of the “perception” that photos don’t count when it comes to piracy, and because there was no possible infrastructure to enforce legal protections.


So, yes, the photo industry has been starved to near extinction, compared to what it could be if it similar legal representation that the music and movie industries do.


“My gosh, I’m getting sad. But I still want to share photos online.”


Don’t misunderstand me; I’m cognizant and sympathetic to the non-professional side of photography and the social value of sharing images, both culturally and economically – including to those photo-sharing sites like Pinterest. There’s no question that people should be able to share images online with others in an unfettered manner that Pinterest provides, as well as every social network.


But to do so in compliance with copyright law would require a series of rights access that cannot be scaled up to serve the public at large without a centralized (and streamlined) rights clearinghouse. Legally speaking, Pinterest should obtain rights from “everyone,” but it’s not possible because people are uploading other people’s photos. If there were a central clearing house open to everyone – say, like the music labels have – Pinterest could enter into a unified license agreement.


Without such a clearing house, the law is the law, and the courts will eventually be forced to reconcile the law with society’s desires. Well, provided cases are brought to court to press the issue.


This is not new. Copyright itself has been a controversial topic for society (and justice) for years, and continues to this day. On one hand, there are many who believe that copyright protection should be lifted, if not severely curtailed, largely in order to avoid this very problem of the social benefit from photo-sharing. Economists, on the other hand, understand that the creative economy only exists because people can earn a living from their efforts—that "human creativity is the ultimate economic resource." (Florida 2002) If they couldn’t economically benefit from their creations, society would suffer more, since the lack of incentives (and hence, resources) would starve an important and socially valuable industry.


The only legal basis for dealing with this dispute continues to reside in the Copyright Act in 1976, which states that “copyright protection extends to original works of authorship fixed in any tangible medium of expression,” including photography, of course. Tightly coupled with the Copyright Act is The Berne Convention, which states that “Copyright must be automatic; it is prohibited to require formal registration.” Yes, the USA provides added protection that permits authors to register their works with the Copyright Office, which then affords them “statutory damages” in legal claims, which guarantees the copyright holder a minimum of $750 per claim, and up to $150,000 if the infringing party “willfully infringed” (that is, with intent). But this registration is not required in order for the copyright to be held by the person holding the camera, and that ownership comes with rights.


So, Pinterest and other social networks are technically contributing to copyright violation by permitting other users to upload unauthorized copyrighted works. This is called “contributory infringement.” This Wikipedia excerpt explains, “indirect infringement arises when a party materially contributes to, facilitates, induces, or is otherwise responsible for directly infringing acts carried out by another party.”


These underlying legal principles of copyright law are subtle, and few are as educated on it as they like to believe—especially corporate law firms that write the legal mumbo jumbo in “terms of service” agreements (TOS). To wit, Pinterest’s own TOS stipulates that when you upload a photo to Pinterest, you are granting it a "perpetual, irrevocable, royalty-free license to use” your photos on its site and "application or services." While this is applicable if you own the photos you upload, you cannot grant this permission for photos that aren’t yours. That is, you are not the legal authority of someone else’s photos. So, Pinterest’s own TOS is unenforceable on photos that the user doesn’t own, which is pretty much all of them. So, strictly speaking, their TOS is toothless, pointless and moot.


But this is also besides the point: the user violated the copyright, not Pinterest.


So again, who’s to complain? To whom? Against Whom?


One could try to sue Pinterest, which is where their lawyers would quickly seek protection under the Digital Millennium Copyright Act (DMCA), which states that websites that host content uploaded by users cannot be held liable for copyright infringement, so long as the site complies with “take down notices” from those copyright holders. Here, the original owner of the copyright notifies the company with a “take down notice,” and the company is off the hook—no TOS necessary.


Many companies – Pinterest, included – very effectively use the DMCA as the “get out of jail free” card, effectively keeping their business out of legal danger.


But once again, we come back to the subtleties of the Copyright Act. As stated earlier, Pinterest could be liable for secondary infringement, which would make them ineligible to seek protection from the DMCA. For matters relating to copyright, courts would have to decide on the merits of such claims solely on case law developments.


This brings us to landmark cases, such as Napster and most notably, Grokster, where courts have established a three-point test to determine if a website “induces infringement”: (1) whether the majority of the content uploaded by users is infringed works; (2) whether the site provides tools that can only be used to infringe; and (3) whether the use of the works are (a) for commercial purposes or (b) harms the commercial interests of the copyright holder.


In the case of “majority of content,” this part is pretty self-evident.


In the case of the site providing tools that are “only” used to infringe, Pinterest’s defense would have to be based on a finding by the Supreme Court in “Sony Corp. of America v. Universal City Studios, Inc,” where the court found that, contributory liability cannot be imposed unless the technology lacks substantial non-infringing uses. Flickr, for example, only provides an “upload” button that allows users to upload images from their own hard drive. This provides “substantial non-infringing uses.” Indeed, the content on Flickr has most of its images uploaded by the original photographers themselves. Pinterest, however, cannot demonstrate this: their tool does not permit uploading photos from one’s own computer; in fact, it encourages users to pin photos from other sites.


The third test –commercial profit– also has roots in the legal doctrine of “Vicarious Liability,” where “courts have extended liability to those who profit from infringing activity when an enterprise has the right and ability to prevent the infringement.”


If someone were to go to the effort of showing that Pinterest satisfies all three tests, the company loses its DMCA protections, and must now face the music. This then re-engages copyright law, where the company could be liable for statutory damages if any of the works are registered with the copyright office. (Many pro photographers whose works are generally passed around the most, actually register their works.) Statutory damages mandate a minimum of $750 per infringed work, although a judge can raise the limit of the claim up to $150,000 per infringement if the defendant was deemed to “intentionally infringe.”


One would assume that if a site lost its DMCA protection because it was “inducing infringement,” then a judge would likely also rule that the infringement was “willful.” Hence, the $150,000 per image claim would be a hefty speeding ticket.


“Sounds troubling for Pinterest! Are they in trouble?”


Probably not. And it’s not because they aren’t in violation of copyright law—they are. It’s back to the basic question of “who’s going to sue them?” Unlike music and movie companies that have hoards of lawyers representing their interests and who aggressively shut down websites and file legal claims perpetually, photographers have no one. As individuals, photographers are too unsophisticated to navigate the difficult and expensive litigation process, so it is highly improbable that many will sue. And even if they did, they won’t be able to do so in a critical mass necessary to materially affect the company the way may music labels can. And even if they could, they’d be up against the same free-speech advocates that defended Grokster. This would not be an easy or inexpensive task, and would probably garner a large push-back from society who already regards photos as “different.”


I don’t mean to “pick on” Pinterest, actually. They are but one of many such sites. Polyvore not only satisfies the three-point test of “inducing users to infringe,” but their volume knob goes to 11: They offer even more sophisticated tools to infringe, including software that specifically designed to copy photos from other sites, while also providing no tools to upload users’ own photos, which flies directly into the face of the definition of Contributory Infringement, and satisfies the Supreme Court’s own language on whether the technology has a substantial “non-infringing use.” Worst of all, they are actually selling products, not advertising, which satisfies “Vicarious Liability.”


And their legal problems go beyond just copyright. Users also upload photos of celebrities to adorn the products sold on the site, which could be in violation of publicity laws if there isn’t a model release. (Cameron Diaz’s picture is one of the most popular.)


Polyvore does provide its own photos, which are legitimately licensed -- namely, from the companies selling the products depicted in the pictures. The test is whether the majority of the content uploaded by users are unauthorized. Other factors that appear to implicate their “knowledge of willful infringement” is a statement warning people not to infringe, and the promise they will take down photos if contacted by copyright owners. While one could argue that they are trying to give notice, this is akin to warning labels on cigarette boxes. No one’s fooling anyone here.


There’s no doubt that Polyvore knows its users are infringing, and it’s certainly possible that they are aware that they are also “inducing” infringement, but they are counting on the same two factors that Pinterest is: society accepts copyright infringement of photography, and more importantly, there are no special interest groups that will sue them for “contributory infringement” on behalf of a class of photographers.


“So, as long as society has accepted photography as a non-threatening step-child in the copyright debate, these sites are safe.”


The force is strong in you, young Jedi.


Still, the risk profile could suddenly spike if there were an unintended rise of those who would intend to assert those copyright protections, which could happen if incentives were to suddenly materialize. For example, a SOPA-like legislation.


“Huh? SOPA? Come again?”


Although the Stop Online Piracy Act is dead for now, the music and movie industries are not about to let it go. Something will eventually re-emerge with new and different terms. We’re already seeing a great deal of anti-piracy legislation coming out of Europe, and Congress and others are under a great deal of pressure to do something (probably after the election season).


What needs to be considered is the unintended consequences that might result if they don’t reconcile the incompatibilities between the social aspect of photography and the fact that it’s a copyrighted work. For, whatever law that has the intention of protecting movies and music just might create a financial incentive for new actors to enter the stage and try to represent the interests of the entire class of photographers, professional and otherwise. And the social networks that use photos are far bigger and vulnerable than the usual targets that music and movie studios attack, escalating the size of litigations that could arise.


A poorly drafted SOPA-like law could affect the internet in highly unexpected ways, akin to the sudden and immediate changes we saw in our political system after the Supreme Court’s decision on Citizen’s United.


“So, do you have a better solution?”


Funny you should ask.


I don’t believe one can ever legislate around this problem. There are two economies at play all the time: a legitimate one and an underground pirate economy. The best you can do is create so much incentive for people to participate in the legitimate economy, that the efforts to pirate become less interesting and less profitable, yielding a progressively smaller proportion of that industry’s total economy. Steve Jobs pleaded with the music industry to remove music locking in song files using the argument that people don’t want to infringe, so long as they can get access to what they want at a fair price. When the music industry finally agreed to remove those locks, online music sales spiked. But the music (and film) industries haven’t kept up with cultural and technological trends in how they handle the business side of their industries. They are still trying to solve 21st century problems with 20th century attitudes.


It’s not that I disapprove of litigation – it’s the music and movie industries greatest advantage. The legitimate marketplace exists because music and movie companies have the infrastructure to enforce copyrights; this is the stick that gets people to seek the carrot, benefiting the entire marketplace financially and fairly.


When it comes to photography, there is no infrastructure for enforcing copyrights, so there’s no viable marketplace. I mentioned that there needs to be a central clearing house for photo rights management: My solution to that is here.


________________________________

On Fri, Mar 16, 2012 at 4:15 AM, wrote:

When a person makes a board and posts other people's photos, isn't it just like sharing a link on a blog? When you click on the photo it goes back to the original website that it was found, doesn't it? I don't understand how that constitutes infringement, to me it's like a beefed-up hyperlink. Are you saying that they are literally taking photos and uploading them somewhere?

Literally copying content is one form of copyright infringement, but that's not what we're talking about here. The fact that the content is merely "displayed on a website not authorized by the copyright holder" is technically an infringement.

To illustrate, let's take your text, and change the word "photo" to "movie":
When a person makes a board and posts other
people's
movies, isn't it just like sharing a link on a blog?

In this case, let's say that those "other people" are movie studios. Here, the the public that views this movie is able to see it on a site that has not been authorized to show the movie. The movie studio doesn't care that the movie also happens to link back to their site, or iTunes, or amazon, or anywhere else. The content itself is displayed on another website without authorization. That is an infringement.

Your question of "link" should not be conflated with "text links," which do not display original content. For example, the text link "click here for this photo" is not an infringement because content is not being displayed.

You may say that "photos are different" because photos aren't like movies, or that movie studios charge money, or anything else. Copyright law does not distinguish between media formats or financial intent or even who the owner is. Copyright law is there to allow copyright holders to choose how their content is used.

Many people think that content can be used unless the copyright holder objects, but that is not the case. Technically, the copyright holder must grant consent first.

I realize this would suddenly make everyone aware that every social network in the world is suddenly in gross violation of copyright, and that's why the legal system and the copyright "process" needs to be updated to reconcile this.

Labels: , , , , , , , , , , ,

Friday, February 17, 2012

Selling Stock: it's about search rank, not price

Yesterday, I reposted an article I originally wrote in 2007, discussing the misconception that microstock pricing is what's driving down overall license fees.

I got a few emails that still challenged my assertion, and it appears I haven't emphasized strongly enough the most compelling arguments supporting this thesis.

All of my research supports the premise that the primary cost of licensing images is not the license fee, but the overhead associated with finding and acquiring the right image. The overhead and administration of a project that would involve photo licensing shows that the actual license fee ranks very low on the budget -- hence, low on the buyer's priority list. My 2007 surveys of buyers showed that.

If the person responsible for finding images for a project is paid $60/hr, and this person spends 2-3 more hours looking for a photo just to pay $1 vs. $50, this translates to paying someone $120-180, just to save $50. People who control budgets know that the license fee for photos is negligible to the total cost of production, even at the traditional stock photo rates. The bigger the project, and lower the proportion of the license fee for the image(s).

Those who sell images are dropping their prices because they're looking at their competition, not the buyer. Further, there is absolutely no evidence to show that sites that have lower prices sell more images. There is definitely a perception that there's a correlation, but that's because people are comparing apples to oranges. Getty sales vs iStock sales are not apples-to-apples because the two entities vary dramatically in search engine results (and other important factors). People talk about microstock sites more, and they link to them (in blogs, discussion forums) and the quantity of images on microstock sites is rapidly growing. So naturally, these sites get higher rankings in search results. Search engines don't rank sites because they have lower prices. They rank sites by size (content), links, and a black magic formula that is best described as "dispersion of discussion in and around the net." In short, microstock sites have more content and get more attention. Hence, better rankings, which translates to more traffic, which attracts more photographers to submit images to them, perpetuating the feedback loop.

In my 2007 survey, those who indicated they were aware of--and use microstock sites-- most don't go to them because the prices are lower; it's mostly because those sites ranked higher in search engine results, where the buyer starts.

Because search engine ranking drives traffic -- especially the untapped (and unaware) segment of the global economy that doesn't use stock agencies -- and because the greatest cost in photo acquisition is time, not the license fee, 90% of the time-savings is the image results the user gets on that initial search. If it takes the buyer to a stock agency site -- microstock or otherwise -- then the deal is nearly done. Price notwithstanding.

This is primarily why I have advocated for years that stock sites should focus their entire effort towards optimizing search engine rankings. While they could have done something about it in the past, the rise of social networks and the plethora of image-related websites and apps has made it impossible for agencies to rank highly on image-search rankings on their own. In today's market, they have no choice but to either partner with, or acquire/be-acquired-by a social-networking site.

The Getty<->Flickr combination is a very pragmatic example. Yahoo is circling the drain, and it needs to shed its non-performing assets and focus its attention on ... something. Whatever that is, it isn't Flickr, and there aren't a lot of buyers that would be interested in that asset, except for Getty or Corbis. The combined product would involve retooling Flickr to be far more socially active (to keep up with modern social networking trends), and to integrate licensing/acquisition into the user/social experience. Most importantly, to provide incentive programs for photo submitters to participate economically. (I've written a great deal about this in the past.)

Of course, perhaps Yahoo should just buy Getty. Facebook is getting into the game, which tends to lead one's eyes towards Google, but they are still struggling to play catch up in the social-networking arena, and their photo division is not run by someone with a disposition towards stock or an awareness of the economics of the photo industry. The company is more interested in building assets that support their advertising model. There's no evidence that "licensing" is on their radar--a pity because they would be on the forefront of the Web 3.0 economic model, where images would play a huge role. (See here.)

In the meantime, there's a $25B shadow economy in peer-to-peer photo licensing that's up for grabs. (See here.)

So, you ask, "how do you convince agencies of this?"
I've been trying since 1998.

(For fun, see this web archive of my site from 1999 discussing this topic.)

Labels: , , , , , , , , , , , , , ,

Sunday, November 13, 2011

Creative Commons Effect on Photo Licensing

Julie Bernstein asked me the following question: "I am curious if your views on Creative Commons have changed since the four articles you published on this topic in '08."

Julie is referring to these articles (part1, p2, p3, p4) where I describe the CC as a great licensing method for almost all media types except photography.

In summary, what the CC has done is create a legally legitimate infrastructure for those who freely share copyrighted works. Before CC, such activity was technically an infringement, because the the publication of creative works requires consent of copyright holders. CC clears up that technicality, which is great. But it has inadvertently given people the impression that it has affected the licensing industry's pricing structures.

CC has not affected the greater licensing market (or prices), largely because of risk: CC has no centralized authority to assure that content is either submitted properly or used properly. Because it's so easy to game the system on either side of the photo (the supplier or the user can sue the other by luring them with a legally misleading scenario), the financial liability for anyone with a lot to lose is simply too high, especially given that traditional license fees are so minimal. So, the majority of image buyers simply stay away from CC.

Now, this is not to suggest there's something wrong with the CC model in principle. I'm a big advocate for it in all other contexts. Indeed, it was born out of the "free software" meme that was popular in the 1980s and 90s, when Gnu Public License (GPL) and other models were the precursors to the "open-source" model we still enjoy today. These are great innovations in licensing because they allow intellectual property to be used for the greater good, while also allowing for commercial use of those innovations.

But CC in the world of engineering is entirely different from photography. Engineering takes a considerable amount of time, resources and (usually) teamwork to produce anything of value that those in the open-source community would use. As such, the kind of content there is proportionally minimal, and each work is substantial and recognizable, making infringements quite easy to spot.

None of this is true in photography -- trillions of images are produced daily, it's impossible to track any given photo, or whether it is "legitimate" (either by the owner or the user).

So, sure, in a world of honest people that want to freely share their content in a peaceful corner of the image licensing market, CC is great. The CC market is growing, but the perception is only as a measurement of itself, not the total licensing market. An article on that topic can be found here:
http://www.danheller.com/blog/posts/total-size-of-licensing-market.html

Lastly, it's natural to ask, "If CC is so easy to game, why haven't we seen it?" The answer is because the market is so negligible. Economists often use crime data as a reality check on the economic activity they think they're aware of. The higher the crime rate, the more economic activity there is, and there's usually parity between that activity and the presumed size of a commodity's market. If there's little crime, the market size isn't big enough to warrant the effort. If CC were to genuinely gain momentum, it would attract those who would game the system for profit, which itself would have a cooling effect, bringing its popularity back down.

For the record, I've proposed that the best way to assuage people's risk concerns about CC is to use the "copyright registration" system. The CC foundation should have a submission system where those who want to submit images for CC licensing would bulk register those images to the copyright office. This gives them the right to file claims on behalf of the copyright owner, which is how major stock agencies like Getty work. Registered images are eligible for higher level of copyright protection, and there are federal penalties for fraudulent use. This means that users of CC images can be protected from invalid claims by those trying to game the system because this is built into the copyright act's provisions. Similarly, authors can be assured of CC compliance because non-compliant users could be subject to an infringement claim. Yes, you can sue someone for copyright infringement, even if the license fee were zero, because the infringement is another form of "breach of contract." Here, the user of a CC image agreed to the terms of CC by (for example) citing copyright ownership. Failing to do so is an infringement of that contract, and is therefore subject to the statutes provided by copyright law.

This would not only allow CC to have actual teeth, but the trust would go up as the risk comes down.

But such an infrastructure would be quite expensive to operate. That'd be a tall order just to create a system that brings the license fee for a commodity down only a few dollars, even if it is only to zero.

Labels: , , , , , , , , , , , ,

Monday, February 14, 2011

Search Engine Optimization and The Long Tail

I was inspired by an entertaining article I read in today's New York Times titled, The Dirty Little Secrets of Search, detailing the rise and fall of JC Penney's Google rankings. Turns out, JC Penney's SEO consulting firm allegedly bought a huge number of paid links on websites, most of which aren't actual sites at all, but domain names purchased solely for the purpose of placing links to PC Penney. Google takes this very seriously, and has been known to eliminate sites completely.

The rationale for this approach is, as most people know by now, that your ranking is governed most largely by the number of other sites that link to yours. Unfortunately, what many people still don't know is that gaming the system doesn't work. (Link exchanges are a sure way to lower the ranking of both sites that link to each other. That's why JC Penney's SEO firm just created sites that had one-way links.) While it'd be nice to have organic linking, where people simply "talk about you" (and provide a link) on many websites on the net, that's not so easy to do and takes a lot of time.

In this day and age, if you're going to succeed as a stock photographer, you have no choice but to figure this out. This strategy begins with two questions: 1) which keywords or phrases do you want to rank highly for, and 2) how do you seed yourself around the net?

The answer to the second question begins with the first: find the right keywords.

Here is where most photographers (and agencies) get it wrong: they shoot for keywords like, "stock photography," and other industry trade terms. But this doesn't work so well. Google's Traffic Estimator shows terms like "stock photography" yields only about 90,000 global monthly searches. Sites that rank highly for only a few keywords or phrases never do well, even for popular search terms. Instead, reach for many search terms -- as many as possible.

My site (danheller.com) ranks in the top five positions on 751 search terms, and 1205 search terms rank in the top 10 on Google Search results, according to Google's Webmaster Tools. But I'm not actually trying to rank highly for any given search term at all. That would be futile. Odd as it may sound, I rank #1 for "stock photography business," but I swear I didn't try to. Of course not, because that search term doesn't generate enough traffic to warrant investing any special time or effort. That's the point. This is the "long tail" approach to keyword indexing: it's about breadth, not depth. I don't get that much traffic to any single page. By ranking highly in such a vast number of terms, it's the aggregate that matters.

All this starts with simply being indexed. That is, search engines have to know what words and phrases you have before it can rank them. Choosing the right words is one thing, but you also need Google to trust your keywords. In other words, trust you. Unlike standard text on a page, which Google is good at, photos are different. An algorithm doesn't know what's inside a photo -- it has to look at other characteristics to determine its content, such as surrounding text, the name of the page it's on, and of course, its metadata. In particular, the "keywords" tags embedded in the IPTC header of the image file.

Once again, here's where most photographers and agencies get it wrong: they "pollute" their keyword lists with dozens, if not hundreds, of phrases and expressions, hoping the target image will come up as a search result for any one of them. But Google will actually penalize people try to game the system with "black hat" approaches, like using repetition (singulars and plurals together), lots of synonyms, intended misspellings (by seeing both the misspelled and correctly spelled words together), and tons of generic terms (such as "photo", "image", "photography," etc).

Products like Cradoc's Keyword Harvester and A2Z Keywording each suffer from (and perpetuate) this problem. The main reason is because they are trying to anticipate what a searcher might look for. This is not only impossible, but the mere attempt reduces your credibility index in the eyes of almost all search engines.

Almost all? Which search engines does it actually work for? One of the people responsible for this policy told me "microstock agencies is where our customers submit their photos, and those search engines are not that smart. So, we have to be thorough."

True enough, but this raises two issues. First, despite the fact that microstock websites are popular among amateur photographers and a growing population of desperate pros, looking to pick up the pennies from as many sources as possible, the vast majority of those looking to license images don't go to stock agencies. They go to main search engines.

Second, even among the brain-dead search technology employed by stock agencies (except for Getty's whose search technology is quite good), proper keywording techniques still perform quite well at those places. The reason is that people searching for images don't go about it in the diligent, thoughtful way that photographers think they do. People do not search using conceptual terms that those who sell keywording products would lead you to believe.

Keywording properly is really boring, and far less time-intensive than people make it out to be: just the basic "facts" about the photo can be described in a handful of terms. The search engine will do the hard part. Granted, this is a bit simplified, because it doesn't address issues like word definition ambiguity, synonyms, and so on. But this isn't done by humans anyway; it needs to be handled by the search engine's heuristic engine. True, stock agencies don't have them, but again, the trade off is whether to achieve "good enough" with the less-frequently used stock agency or the "proper" method advocated by the search engines.

This is why the "proper" method achieves the best of both worlds: you will be indexed properly and given higher "credibility" with public search engines like Google, and you won't be penalized by the microstock agencies even though images might only use a handful of keywords, rather than dozens or a hundred.

The next question is how to get all those coveted links from other sites to direct traffic your way. This technique is not easy; it requires work. You need to write a lot, post to discussion forums, socialize and network, be on the "inside" with industry people, and above all, talk about what you know. And here's the real hidden secret, I'm not talking about photography. The discussion forums, industry people and the topics you talk about are best when it's something other than photography because it's highly likely that you're an expert at something other than photography.

Of course, if you are well-informed about photography and are regarded as a leader in the field, then go for it. But if you are, then you're probably not reading this... at least, not with the goal of improving your photography business. I am better known for my business analysis, which happens to be in the photography field, than I am for my photography as an art form. That I sell lots of images (prints and licenses) is not a byproduct of my artistic skills. It's the byproduct of having published so much about the business of photography.

The more you engage in discussions online and offer useful, insightful and meaningful commentary, the more people will link to you. Offer to write for magazines. Try even writing a book or two. Sure, it's an investment of time. What'd you expect? That it'd be easy?

Labels: , , , , , , , , , , ,

Monday, June 28, 2010

Getty and Flickr: Prophesies Coming True?

People have been emailing me copiously, asking for a statement in response to the new relationship between Getty and Flickr, where Flickr members and visitors can work with each other through a new program with Getty Images called “Request to License”. The details of this program are listed here. From that page:

When a prospective licensee sees an image marked for license, they can click on the link and be put in touch with a representative from Getty Images who will help handle details like permissions, releases and pricing. Once reviewed, the Getty Images editors will send you a FlickrMail to request to license your work, either for commercial or editorial usage. The decision to license is always yours.


Why are people asking me about this?

For years, I've been proposing that precisely this model be implemented. Most of my blog entries in 2007 and 2008 articulated this very model. The first was on Feb 13, 2007, in an article titled, "The future of photo sharing sites and agencies". There, I predicted the inevitable convergence between companies like Getty and Flickr:

I believe it will invariably happen that major photo agencies like Getty and Corbis can (and should) move into the consumer market. Consider what would happen if major stock agencies expanded their businesses by opening the flood gates and letting everyone in. By removing the barriers that require photographers to "submit images," and having a separate portion of their sites be entirely open, much like other photo-sharing sites are, they would give more options to buyers, and provide more opportunities (and greater incentive) for photographers to join at all levels. Getty owns iStockPhoto.com, which is a microstock agency that sells images for much less, but this is not a consumer-based, social networking style photo sharing site like flickr is.


The key here is in italics: microstock agencies are not social networking sites, they are therefore limited by both buyers are sellers than the social-networking sites. My premise for this logic is based on my years of research showing that 80% or more of licensed images is peer-to-peer, directly between buyers and photographers, not among agencies. You can read this research in the article, "The Size of the Photo Licensing Market"). The summary of that research is this basic truism: Most buyers find images on non-stock agency websites.

On Feb 18, 2007, I wrote how the photo-sharing and social-networking sites can capitalize on this opportunity in an article titled, "Two-Phased Approach to photo-sharing/licensing model". I said:

Phase One of this business will be where a photo-sharing site merely allows visitors to license images directly from the site. Phase Two will involve the distribution of the same photo assets to other sites, much the same way online ad sales are hosted (or "published") on other websites. ... For the sake of discussion, I'm going to assume that the approach ultimately adopted is the one I've suggested in the past: make it pure and simple by giving the user a toggle for setting whether his photos are (or aren't) permitted to be "sold".


And that's exactly what Getty and Flickr are doing now. Over four years later.

You may note that I said there was a two-phased approach. That second model will eventually become part of more photo-licensing business models. (In fact, it already exists, but among companies too small to get anyone's attention--partly because the technology and business models they've adopted do not properly understand and implement the true nature of photo licensing, copyright issues, and potential target markets. This is an aside for the moment; it may come up again when larger players eventually begin to consider the opportunities.)

Speaking of predictions, I remain steadfast in my opinion of the inevitability of what happens next:

In July, 2007, my blog post titled, "The Solution to Getty's Woes" explained how Getty can get out of its financial troubles by simply buying Flickr directly from Yahoo and using it as the main stock licensing engine. The article got into exceedingly detailed analysis of Getty's financial model (and troubles) combined with the explosion of available imagery on sites like Flickr that make this solution not only obvious, but inevitable.

On a directly related note, I called into question the life expectancy of the Creative Commons in this article (2008), where I again proposed that Flickr allow users the option of choosing between allowing their images available for free via CC, or to get income from their images. I said,

...it begs the question about whether enough people would choose the option to "make my images free"(CC) if it were next to the checkbox that says, "pay me a quarter if someone's dumb enough to buy it."

And then there's the buyer. If they were given the choice between "free images, with disclaimers and risks" and modestly priced images without such risks, it wouldn't be very likely that the "free" versions would be chosen very often.

The concept of CC would never survive under these two conditions.


Without getting too far afield, I have no qualms with the CC, per se. It's more about how simplistically it's been designed and deployed. It's just not sustainable in the real world business market. The problem is not the "license terms" and the structure of the legal contracts--those are all just fine. It's the fact that the system can be gamed so easily by both buyers and sellers, that it's too unreliable to be sustainable beyond a small handful of casual users (by comparison to the larger market of stock imagery). The true protections for both buyers and sellers is to leverage the copyright registration mechanism. That is, creative commons images that are also registered with the copyright office lowers the risk both both buyers and sellers, as explained in that article. Since no one is building copyright registration into their online business models, and the CC itself has a fundamental objection to the concept of copyright in the first place, the CC will be relegated to an historical footnote , bringing strength back to the for-fee licensing model. And which brings us back to why I'd always argued that Flickr should have enabled image licensing.

So, why is this all good for the photo licensing industry? I articulate this answer in the blog entry I wrote on March 15, 2007 in the article titled, "Photo-sharing-licensing sites leveling the playing field."

As more companies engage in the business of licensing images, photographers with credibility will gravitate to the sites that offer a better return on their money... In a way, this is how photo agencies started in the very beginning, only better: because photographers don't have to be "accepted," the playing field is much more level, and the market forces can be more free to let the money flow to those who really do merit the higher earnings (rather than at the whim of photo editors). The buyer, it turns out, is the best photo editor, and it will be pretty clear in short order which sites are hosting good, honest content.


I summarize with another excerpt from that article:

...the most basic, fundamental truism about photography remains: there are more people who have it as a hobby than as a profession, and the barrier to entry is low... the honeymoon period for Getty will end once photo-sharing sites become new outlets for photographers where the open market can decide their rates."

Labels: , , , , , , , , , , , , ,

Wednesday, March 24, 2010

2009 Year in Review: Web Optimization

In this second segment of my series, "2009: Year in Review," I discuss issues related to managing my web presence. Some of these methods directly result in income, such as advertising dollars, whereas others indirectly affect income, such my ranking in search engines or by directing traffic towards monetizable content. Nothing discussed here addresses my actual sales and licensing methods, which was addressed in Part 1 of this series.

Web Traffic and Advertising

Traffic to my site has marginally increased by 16% from the same time last year (2008). More specifically, I averaged about 15,000 visitors a day in 2009, but the number would have been much higher had it not been for a technical mis-decision I made during the summer months that dramatically dropped my rankings, which had to do with "keyword stuffing", discussed later. Normalizing for that, my traffic has been pretty steady at around 16-18K unique visitors a day, compared to 14-15K/day in 2008. (Stats can be seen here.)

While that may sound impressive, it's not that simple. There are a number of devils in the details, and sifting through the data is only half the battle. For example, the bounce rate (the rate at which people leave my site after viewing the first page) rose to 8.5%, and the average time on site dropped by 11%. In other words, people are leaving my site sooner than before.

One would think that this is a bad thing, but there's other data that suggests otherwise. For example, advertising revenue more than doubled; in some cases (some pages and topics) tripled and quadrupled. All those people "bouncing" away without spending time on my site are clicking on ads. For 2009, advertising revenue jumped to represent 17% of total income.

One might say that I'm losing potential buyers to advertisers, but that's not what's going on. Most of the ads on my site are not for photography prints or licensing, which is the lion's share of my online transactions. That is, people are clicking on ads because they decidedly do not want anything I have to offer. I don't care that they leave; it just so happens that they're paying me a effective "exit tax." Or rather, the people who are getting my traffic are paying that tax.

Indeed, this turns out to be mutually beneficial: advertisers whose own sites don't rank well for some search terms, actually get a lot more relevant traffic from my site than they would if they paid to get onto Google's search page directly. That is, they'll pay ten cents to a dollar per click to put an ad on my page (through Google's adwords program), compared to twice or three times that much to put the same ad on Google's search results page. They may not quite get the same number of total traffic, but they'll get much more relevant traffic that converts to revenue if they place those ads on my site (or any of the other top-ranked sites). This kind of advertising-indirection costs them less, they get better bang for the buck. Best of all, I get a cut of it. :-)

I should point out that this isn't always so straightforward for advertisers, because targeting a specific site can be costly (in the form of lost opportunity, not necessarily money) if that site isn't consistently well-ranked. That is, if they target a site that appears to rank well sporadically (because their content changes), they could get a boost of traffic for a short time, and then go dark. Since my site has been around for a long time and is generally stable, this risk is not a concern.

In fact, many advertisers come directly to me and pay me to put their ads on my pages, rather than going through Google. There are advertising aggregators that have clients that pay them to do this analysis, and my site is coming up more often in their radar. My advertising rates are not based on clicks or impressions; they're flat fee rates, which advertisers like a lot for a high-traffic site like mine.

This then begs the question: what was the actual end-user looking for that they landed on my site, even though I didn't have what they were looking for? Why am I ranked so highly for them? Isn't that a problem with the search results?

First of all, the bounce rates are still quite low. Google does accurately put users on pages that match their searches. Of the low number of people who bounce, it's usually because they used the wrong search terms in the first place, and Google couldn't possibly know that ahead of time.

Take the Olympics in Vancouver, for example. If you search for "photos of vancouver", I'm currently ranked #8 on Google. (Before the Olympics, I was ranked among the top three.) So, I get a lot of people looking for olympics photos, even though they didn't use the term, "olympics" in their search query. When they don't see such images on my Vancouver page, users click on an ad that gets them where they wanted to go.

Vancouver is only one of a long list of examples. At the moment, I score very highly for phrases like:

  1. "black and white pictures" (Google Rank: #4)
  2. "what kind of camera should I buy" (#6),
  3. "learning photography" (#2)
  4. "photography business" (#1)
  5. "model release" (#1)
  6. "star trails" (#1)
  7. "fill flash" (#1)
  8. "photographing people" (#1)
  9. "selling prints" (#1)
  10. "photography marketing" (#3)
  11. "sahara desert" (#5)
  12. "stairs" (#6)
  13. "photos of doors" (#1)
  14. "photos of new york city" (#3)
  15. "photos of san francisco" (#1)
  16. "photos of kids" (#1)
  17. "photos of united states" (#1)
  18. "photos of patagonia" (#3)
  19. "photos of cuba" (#1)


These are but a few among hundreds of phrases that Google ranks my site and/or pages among the top-five. But the key is that these terms are generic and they themselves do not bring traffic that can be attributed to a single dime of sales revenue.

While they are good for generating advertising revenue, there's an even better benefit to ranking high for generic search patterns: Non-buyer traffic out-strips buyers by orders of magnitude, and any traffic--buyers or not--contributes to the overall ranking of my site. When people search using more specific terms (for content that they do want to purchase), my site will rise in those search results, yielding sales.

So the objective is to have as many pages rank as highly as possible. One key strategy here is that I don't particularly care to rank highly for any single or small set of search terms--that doesn't necessarily benefit me. It's just having my site itself be indexed well for whatever content the search engines deem appropriate. And therein lies the question: how do they determine what search terms should send users to my site? Since they cannot determine what's inside of a photo the way a human eye does, search engines look for other clues to determine the content of a page that otherwise has very little text: metadata.

Keywording

I've blogged before about keywording; it's a huge topic. I'm not going to reiterate points I already made, but to appreciate how and why I employ my keywording methods, you need to at least understand this very basic set of truisms:

  1. Most image buyers use search engines first, stock agencies second.
    Search engines act like "metasearch" for all the stock sites, as well as many other image sources, including mine, yours, everyone else's. It's best to use keywording techniques advised by search engines, not stock photo agencies.
  2. Search engines are intelligent about search queries.
    Unlike days long ago, they know all the synonyms that are related to a common root. So, you do not need to include the singular and plurals, all the variants of "dog" (canine, puppy, pooch, etc.), and so on. What's more, intelligent search is becoming more common, even among stock agencies. The need to stuff your images with synonyms and other related keywords to make your list "more thorough or complete" is gone. In fact, attempting to do so can backfire on you. (More about that later.)
  3. Controlled Vocabularies are a complete waste of time.
    There was once a time when such lists were useful, because it made the job of image search much easier for unsophisticated (brute force) search algorithms. Controlled vocabularies helped you use a small, consistent set of words, which kept you from using dozens of similar words that might come up with different search results when the user input search queries.

    While that premise was useful, it only addresses half the equation: the weakest link in search is not you, it's the end-users. Or rather, the search queries they submit. These people are not going to conform to controlled vocabularies. So, in order to map their queries to your images, their input text has to be converted to root words anyway. If the search algorithm is going to do this to end-user queries, it can (and should) also do it with your keyword list. Forcing you to conform to a list becomes a waste of time.
  4. Keywording should take only a few minutes and minimal thought.
    It's very easy to over-think how people might find your images, or to worry that your images might not be found if someone uses a series of queries that you didn't think of. But this kind of over-thinking can negatively affect if and how your images are found. End-users learn very quickly to be very conservative in their search queries, or they will get a lot of irrelevant results, rapidly wasting their time. They may experiment with creative, conceptual, or "refined" queries to see what they get, but it doesn't take long to learn to "keep it simple." So should you. Keywords should include only the most basic, obvious, and prominent items in the photo. Search engines also rank the quality of photos (and the sites that host them) on their brevity. More than ten keywords will diminish a photo's rank because it usually means that someone is going to stuff the keyword list with unrelated words in an attempt to game the system. This is a common technique among photographers who submit their images to dozens of microstock agencies who do not enforce such restrictions, and who use brute-force (letter-for-letter) search algorithms. Keyword stuffing--also known as "keyword pollution"--has proven to be effective for such photo sites because it allows those images to be found ahead of other, potentially more relevant results for any given search.


In fact, I fell victim to "keyword stuffing" myself midway through 2009. In my automated keyword algorithms, which normally strips redundant or "similar" keywords, I had thought I was being clever by adding in location information (city, state, country) into the keyword list. Yet, what I found was that because the IPTC data already had these keywords, which search engines tap into, and because my keyword list grew (unnecessarily) by three more words, this dropped my rankings down by several notches, which kept me out of the "top fold" of search engine results. It's a huge deal dropping from #3 to #6 or #7 for a given search term, and you can see the results of this in my site traffic data over the summer of 2009.

Needless to say, this cost me quite a bit in traffic, which affected every other aspect of my business, from sales to advertising rates.

You can imagine, therefore, that "effective keywording" (so that images and website are deemed "credible" and ranked highly) is a hotly debated issue in the photo community. It's also one where entrepreneurs try to come up with solutions--some good, some not so much.

One example is a product "imense annotator" (annotator.imense.com), which has some interesting ideas, such as an image-recognition algorithm that tries to guess keywords that might describe the people in an image. It will do a reasonable job in ascertaining the ages, sex and ethnicity of people in a photo, and then attach those keywords to your images. Clever, and possibly quite useful more to a stock agency than an individual. This is because agencies have millions of images to process, none of which have been (or will be) seen by company staff. On the other hand, original photographers that shot the images could do this task quite easily on their own. One can only shoot so many images in a day, and since one has to eventually go through a manual (if not minimal) keywording phase anyway, one can assign the keywords associated with the "people" photos as part of that process. This shouldn't be all that time-consuming for reasonably well-disciplined photographers. And human analysis on such things is always going to outperform a computer. (Yes, I say this as an active programmer.)

(Note: The annotator only does people/facial recognition.)

All other aspects of annotator look and sound cool, but are considerably less effective in practicality. Again, these include "commercial vocabularies", "crowdsourcing" and "controlled vocabularies." As noted earlier, these ultimately contribute to the perils of keyword stuffing that search engines don't like--and which only serve to confuse stock agencies' less sophisticated search algorithms.

Another thing to keep in mind is keywording is often done once, and then you never touch those particular images again. Therefore, whatever you use as keywords today are likely to stick with your images long into the future. But technology doesn't sit still--especially image-recognition and search algorithms. For these, time has a tendency to speed by rather quickly. Before you know it, most search engines will be incorporating the same sort of algorithms like the annotator above. In fact, Google's own image recognition features are rather well developed, and can be seen in action if you use their Picasa image management solutions.

In any event, the point is that keywording is a classic case where "less is more." Images should have minimal base tokens in the keyword list; the search "intermediary" interprets the uncontrolled end-user queries and maps them to the minimal keyword list in your images. This is and will always be the most effective way for images to be found.

While I don't necessarily fault software companies for coming up with creative ways to "enhance" keywording, I draw the line when companies actually recommend methods and behaviors that are wholly counter-productive. An example is Cradoc Software's latest product, fotoKeyword Harvester, a product that does a form of semi-automation of keywording your images. While I am a fan of the company in many ways because it tries to also be the photographer's "coach" on many vital business matters, it has never been on the forefront of the photo business--rather, they seem to be stuck in the 1990s with many of them. Alas, most of their advice, while applicable 10-15 years ago, is well behind the times today.

In the case of the Keyword Harvester, the company sent out an article titled, "best ways to keyword images using concepts and attributes." A quote is: "You'll need to start paying attention to how images convey messages in advertising." They say:

One of the most valuable types of keywords for an image are things called Concepts. A concept is a term that describes non-concrete aspects of your image, an abstract idea. Concepts are used by advertisers to sell their product with the use of your image. They want the consumer to think of something specific when their product is thought of. (...) For example: Wells Fargo Bank uses images of cowboys, wagon trains, horses, and the wild west to promote their business. The concepts for these images are: excitement, freedom, trust, historic, strong, powerful.


There are several problems with all this. First is one I highlighted above in my bullet list: photo searchers (commercial or not) do not use conceptual search terms very often--at least, not with much success as they once did when the stock industry was far smaller, before digital images, and before the internet--a time when almost all stock sales were dominated by Getty Images. Back then, yes, conceptual keywords worked. And this was because Getty internally controlled all keywords for all images. Also, they had their own intelligent search, and they controlled the images in their databank.

Today, images are found in many places, are keyworded by arbitrary staff--or worse, photographers--and the consistency is impossible to centralize and manage. The direct result is that photo buyers don't search the way they once did. (This is an example of Cradoc seems to be stuck in the 1990s.)

It's easy to put this to the test: go to images.google.com and search for the "conceptual keywords" that Cradoc said represented the kind of themes Wells Fargo uses in their imagery. I tried every word on their list, as individual search terms, in pairs, in triplets, and as the entire group. Not one single set of results from these queries contained images that would ever be used by Wells Fargo. They are totally unrelated to all their business models. This is not unique; it's rarely ever the case that conceptual keyword searches yield desirable results. That's why most searchers don't use them anymore.

By contrast, if you search for images based on the actual elements used by Wells Fargo imagery -- cowboys, wagon trains, horses -- image search results show many images similar to those the bank actually uses.

Again, the lesson: keep it simple. Don't get clever. Do not try to anticipate what the searcher might use as search terms. Photo researchers are more afraid of you than you are of them. They are going to keep it simple, too.

I can verify this with my own statistics: My site gets about 19,000 search queries a day on my own search pages. Of the search terms I get, 99% are for very specific items. Furthermore, when someone actually licenses an image from me, and I track their search patterns that lead up to the sale, it is never the case that people use conceptual terms.

In preparation for this article, I interviewed one particular client about how he tends to search for images. He said, "I found that sites are so inconsistent about search terms, that I've learned not to use big words. Just be as specific as possible to the actual things I want to see in a photo."

When I asked him how he chose the particular photo he licensed from me, and what search terms he used leading up to it, he said he wanted a "futuristic landscape." When he tried that phrase (and derivatives, such as "future" and "cityscape") on Google, Getty and Corbis, he got nothing like what he wanted. So, he just got specific: "glowing buildings", which lead him to the image he licensed from my site, which can be seen here.

Keywording Methods

So, let's get to brass tacks: how should you keyword your images? Google has a document called, Google's Search Engine Optimization Starter Guide, which includes tips on optimizing your images for search. It all boils down to:

  1. The image's filename should include the most relevant elements of the image.
    For example, if it's a photo of a boy and a dog, use "boy-dog.jpg". If you have many such images, use sequences: boy-dog-1.jpg, boy-dog-2.jpg, etc.
  2. Use keywords sparsely.
    The more keywords you try to associate with an image, the more you dilute it, bringing down its "rank" and relevancy (and credibility) with search engines, or with given search queries. This is because search engines use two key metrics to determine how well a given image matches a search parameter: the ratio of matches between an image's keyword list and that of the search query, and the filename of the image. For example, if the user entered the query, "boy and dog", the search engine sees two words: "boy" and "dog." (It throws out filler words like "and.") Here, the image named, boy-dog.jpg has a 100% hit ratio of query terms with keyword terms, and the keywords were in the filename. Note that the actual photo itself may very well be that of a fish and a boat. (Google doesn't actually look at that, because, well, it doesn't know how.)
  3. Avoid using synonyms and other "related" terms in keyword lists
    That is, do not attempt to be thorough in describing images with keywords. That's not your job. Search engines already know how to do that. They've got thousands of programmers with PhDs doing that for you (and for the end-user). The more you try to "help," the more you're actually interfering with the process, which reduces your relevancy and ranking.


The good news about keywording is that proper and effective use of keywords is extremely simple and shouldn't require much (if any) thought or time. Using myself as an example, my workflow involves two phases: the edit phase (where I rename all my photos so that their filenames reflect their content), and the keywording phase, where I apply individual words to images--usually in very large batches.

For example, let's say I'm on a photo shoot of a boy and a dog. After editing out the stuff that gets tossed, I'm left with several hundred images, where I then name them just as recommended by Google: boy-dog-lake.jpg, boy-dog-bridge.jpg, boy-dog-1.jpg, etc. In order to assure the highest ratio of search queries to keyword terms, I try to limit filenames to two to six words, though most are either three or four. This is a difficult decision because if I use too many words, I may "match" more queries, but the ratio will be diluted. If I use too few words, I will rank highly for very narrow searches, but may miss more opportunities. This trade-off is a zero-sum game, so rather than try to game the system, I just be honest: determine what's in the photo, and use that as the filename.

Any words that may be "in" the photo, but seems to be less relevant are then added to the keywords list in the image's metadata. And even then, I rarely add more than two or three words, usually modifiers such as "young" or "funny."

Naming files is often very quick because most are batches of similar images. One only needs to browse a given gallery on my site to see the number of similar images that are shot together. The keywording process is similarly fast, also involving mass-assignment of specific, unambiguous words to large batches of images. My rule of thumb is that keywording thousands of images should take no more than 30 minutes.

Most any image-management software can add keywords; I happen to use Adobe Bridge, which is bundled for free with Photoshop or any of the creative suite products.

Note that if you inspect the images on my site, you may notice that they appear to have lots of keywords. Most of these keywords aren't actually in the images that I process--these are added later by an automated post-production algorithm that generates all my static html pages. I do all this to present hints to the end-user for suggested related search terms to stimulate new search ideas.

Maps

The newest addendum to my website is the use of Google Maps. Essentially, each of my web pages incorporates a google map to represent where every photo was taken. While it may seem frivolous, there's been great advantage to the maps. (It also wasn't entirely easy; Google set up the whole mechanism for the sole purpose of presenting maps based on specific street and/or mailing addresses. I have no interest in that level of detail; I just wanted to generate maps for generic locations, like city/state/country. Well, that isn't quite so easy because there are many streets named after cities, states and countries, and there's no way to tell Google maps that I'm not interested in street addresses, just general city maps.)

Though I instituted maps onto my site late in December, the effect its had on my traffic and ranking has been a surprise. Search engines seem to give extra boost to web pages that are geo-tagged--that is, they indicate location. When people search for images where the search parameters include a location, my pages get an additional bump. I've seen about a 10% boost in traffic two months after having introduced geo-tagging onto my web pages, and I look forward to seeing more data to quantify the extent to which geo-tagging has long-term benefits.

Labels: , , , , , , , , , ,

Sunday, November 29, 2009

Why there's no one-stop shop for photo buyers

I got an interesting email today from someone that prompted me to address a question on many people's minds: why hasn't a single website emerged as the "primary" place to license images? As those in the photo industry know, Getty sells quite a bit, and microstocks fall behind them, followed in turn by a smattering of pro photographers and others who do well as individuals. But, with the trillions of photos on the web, and with the sheer magnitude of opportunity, what's really the barrier to growth?

Here's an excerpt from his email:

do you think that there could be an issue with just straight up too much content for buyers (all buyers)? Say for example - there was a website that everyone knew was the place to buy and sell images for any sort of use (commercial, etc.)? Couldn't there still be too many shots of a 'dog'? ... if all the buyers/sellers universally knew this was the place for images - wouldn't the back-end functionality of a site like this take an army of programmers to design? I have to wonder why something like this doesn't already exist or is in development by a major player like google?


Saying that there's too many photos out there is like saying that Google can't index the web because there are too many sites. And thinking that there could be a single go-to site for buyers and sellers is like saying that there's could be single go-to site to buy electronics. There aren't because there's competition, etc. But, the reason why there's a viable, stable market for electronics (unlike photos) is because there are mechanisms in place that help establish price points, distributors, manufacturers, and so on. In short, it's a mature industry.

The same cannot be said of the photo industry for a variety of reasons.

To begin, it's not that there's "too much content" it's that there's no reliable mechanism for sifting through it. Go to images.google.com and type "two men shaking hands" -- a common image search for business purposes. Though the matches are generally accurate, the results are also entirely arbitrary. We have no idea if these are popular images, or they are shot by famous people, or if they are "current", or even whether the source (website) is ranked highly.

The same is true for every website that displays photos--buyer website or not. "Arbitrary results" is why buyers have trouble finding what they want quickly and easily. Yet, despite the huge number of sites that talk about "swine flu", a quick search on that topic usually gets you exactly what you want on the first page. You can even misspell it -- say, "swing flu" and still do well.

So the first problem that people have to solve is search. And this is irrespective of where people go. Now, people can argue about whether real buyers go to search engines or to stock agency sites, but the technology barrier exists nonetheless. Who's going to solve it? It cannot--by definition--be a stock agency. Why? Because they will not (and cannot) return results that are photos they cannot represent.

Whether now or in the future, search engines will be how most people (yes, buyers) find photos.

Now that's not to say that companies like Getty couldn't solve the problem. In fact, they should. To do so, however, they would change their business model from being an exclusive seller to one that acts as a proxy for others -- a grand middle-man, much like how Visa and Mastercard merely enable transactions with an infrastructure. They don't actually participate in the transaction itself.

The reason why Getty won't be coming to the table here is that they suffer from two basic errors in their understanding of the photo industry: first and foremost, they don't see the market outside of their existing world of traditional ad/media buyers. Yes, that's a big industry, and Getty services them adequately. But it's tiny in the broader world of image licensing transactions. And this leads to their second grave misunderstanding: the belief that the best way to service buyers is to have limited content that's hand-edited by seasoned photo editors.

This is not in keeping with the internet today. As we have all learned in the past decade, Web 2.0 means that the "crowd" is the editor, and order among the crowd is achieved by applying intelligent ranking algorithms, tracking their behaviors, and mining user preferences to glean predictability. Getty doesn't do this. No one does this. (Well, google does with their regular text search for non-image content.)

Getty's model of knowing and understanding what buyers want is fine if you have one-to-one relationships with them, and that's what Getty's good at. I'm not suggesting that Getty doesn't keep what they have in-house. It's what they don't have -- what needs to be built -- that they lack, and what the industry needs.

For the industry to become "mature", there must exist two main functions: (1) a search and ranking system that returns reliably accurate results beyond any measure we see today, and (2) a predictable and viable pricing model that represents a true market-maker commodity market.

Granted, these are not simple problems to solve, but there is precedent for similar algorithms. For instance, Google took quite a few years before they came up with just the right mix of variables and weightings to determine which web pages match a given user's search criteria. A similar list of factors can be derived to determine quality image searches in a variety of contexts.

In the photo world, everyone looks to metadata (such as keywords) as the primary factor, but that's quickly morphing into something else. It's a longer discussion to have, but it sums up this way: google stopped looking at web pages' "keywords" field in metadata because it became unreliable--the system can be gamed, and bad players were ruining the reliability of google results by lying about the content of their web pages. Google solved this problem by no longer looking at web pages' keywords settings, and did their own semantic analysis of web pages to determine what keywords should really be associated with them.

It's not that "keywords" themselves will go away in images, but the future of image search will go well beyond user-editable metadata. Factors like "age" (current-ness), supplier, number of views, frequency of (published) use, and automated programmatic analysis about images to assess various conceptual attributes as well. Hard? You betcha. Yet, just as it took google years to come up with their "hundreds of variables" that determine a web site's ranking for any given search term, so too will effort have to go into determining the relevancy of any given photo search.

There are various image recognition algorithms that can begin to move in this direction, but once you get into the science of it, the field is much broader than people think. There are proximity algorithms to determine if two images are the same, there are content algorithms to determine image characteristics (color, textures, emptiness, focus, etc.), and pattern recognition (faces, emotions, objects, and other patterns). Underneath every image recognition algorithm must be a hierarchical database of seed information to spin the world into motion so that trillions of images can be automatically processed from web crawlers continuously.

And this science isn't just to determine search relevancy: it would also be used by an auction-based system to set pricing for a trading system, exactly like how google prices keywords for its advertising system. This is not simple college algebra, but it's also not science fiction. Similar models have been built to create efficiencies for more complex markets than photo licensing. (Personally, I envision a structure similar to the formulas used for pricing stock options. Here, rather than having a big/ask market of quotes, prices are derived from external data based on aggregating a weighting of historical pricing patterns for images with similar characteristics.)

So, who's going to do it? Therein lies the $64,000 question.

Here, there are two problems: As just described, there's a lot of technology here, ranging from all the various image-search algorithms to the pricing analytics. This by itself is already way beyond what any one company has done. So, whoever's going to attempt this must likely be large enough to go on a bit of a buying spree.

Secondly, there needs to be a general realization that there's money to be made. This is difficult when the cultural behaviors around photos is to share, steal, or do whatever you want. Most investors don't really see the "vision" of a viable worldwide market with these sorts of things become further embedded in our online social fabric. This, despite the fact that research shows that photo licensing is still a $25B industry anyway. Comprised mostly of peer-to-peer transactions, it is now precisely how the online advertising industry was before Google.

More reading on the true size of the photo industry can be found here.

Google could solve both problems outlined above, but that would introduce another problem: They are in the advertising business. If they were to enter a business that monetized content on the web, it would bring into question the objectivity of their search/ranking algorithms, which is the only reason people trust their advertising model. That is, buyers and advertisers believe the pricing model because Google currently has no financial interest in monetizing the content on any given site.

What the world needs is a search engine from a company whose business model is not to sell advertising. Yahoo is becoming a much more likely candidate for something like this (as I'd noted in this blog post), but the catch-22 here is that if they were that visionary, they'd have started this kind of development with their Flickr property years ago. Yes, Flickr could be a good launch pad for such an endeavor, but Flickr's too busy with other things to bother with such fantasies.

Then there's Microsoft: they have the money and the infrastructure to actually accomplish this feat, especially given their efforts to build a good search engine. Their physical and political proximity to Corbis (started by Bill Gates) could also be a substantial contributor to the transaction and pricing models that will be needed.

But I dont' have faith that we'll be seeing headlines in these areas anytime soon. That may change with the evolution of Web 3.0... For that, see this post:

http://www.danheller.com/blog/posts/economics-of-migrating-from-web-20-to-30.html

Labels: , , , , , , , , ,

Friday, October 16, 2009

Might Picscout Ultimately Cause Yahoo to Acquire Getty?

I realize the title of this blog is rather provocative. But let me lead you through this.

It all starts with David Sanger's blog on picscout's new Image Registry and Image Exchange, which is the system that Picscout uses to index images and bring buyers and sellers together through third-party licensors. David makes insightful comments on three critical points.

First, his point #2:
Picscout aims to take a percent of sales, noting on their site: “ImageExchange acts as an online affiliate program, sharing image-licensing income between PicScout and licensors.” This will reduce the percent that goes to the photographer.


David is not the first to observe this, but it illustrates how the big picture is being missed. The premise begins with the fact that the universe of images users (some of whom are active buyers, but most of whom are not) use applications that produce documents (digital and print). Those applications are developed by third party Independent Software Vendors (ISVs), such as Adobe or Microsoft. If the applications that ISVs produce adopt the Picscout API to hook into the registry to identify images the user is using in his document, those users will not only be automatically notified they are using copyrighted images, but will also be given the opportunity to license them. This concept isn't far-fetched--exactly the same thing is done when users try to view movies or listen to songs on some devices.

However, because such a thing is not yet done for images, it has the potential to transform the stock licensing industry. If enough ISVs adopt the API and hook into the registry, a critical mass of users will be invariably recruited into the photo licensing economy. The more ISVs that adopt this API, the more applications will be using them, which casts a wider and wider net of users... who themselves become image buyers.

Here's the hitch: those ISVs will not adopt the API unless they have a stake in the game. That is, a cut of the license revenue. Unless someone has another carrot to wave in front of those ISVs, that's the only way to get them to participate in the program. If ISVs don't adopt the API, this whole discussion is moot. No one uses the registry. Game Over.

Therefore, the game is to capture the ISVs. And the only financial incentive they can possibly have is to participate in the licensing model--that is, a rev-share. This has the even greater advantage of giving the ISV even more incentive to get their own users to license images. The more they license, the more money the ISV makes. The ISVs will not just promote these features, but they may make it pretty darn difficult for users to avoid these features.

Imagine what Adobe would do if they had the ability to get a cut of a $10B economy if they just added a feature into InDesign that assured that the photos being used in any given document was properly licensed.... much the same way an iPod assures that the movie it's about to play has been purchased.

This is the same model I've described in my article, The Economics of Migrating from Web 2.0 to Web 3.0: convert the vast majority of image users into image buyers, and sales volumes go way up.

So, that David observes that photographers' percentage of royalty goes down is a true statement, but one that clearly misses the big picture. Obviously the ISV rev-sharing cuts the pie into smaller slices, but a smaller slice of a much larger pie.

David then makes another keen observation in point #7 about Picscout's underlying technology:
Evaluating an entire page of thumbnails is time-consuming. Each thumbnail must be downloaded and analyzed by the PicScout servers before returning index comparison results...


Though David only cites the Google search as an example of how users expect "speed," this is only the tip of the iceberg. Picscout's web browser plug-in that examines google searches is merely a prototype to demonstrate how the API works. Once again, the real goal is to capture ISVs.

But David's observation is more prescient than he may have thought, for performance is probably even more important than rev-sharing by ISVs. If their apps degrade in performance by using the Picscout API, they won't use it, irrespective of rev-share.

The technology Picscout has introduced is clearly first-stage prototypes to introduce the business model and be the first on the map. Yet, it's also Picscout's Achilles Heel, as there is a race about to ensue.

Let's not be naive: Picscout is not the only company on this track. Image-recognition is a science that's akin to text search: there are many ways to do it--some better than others--but it only needs to perform to minimal threshold for the business model to succeed. Many other factors dictate success or failure. Sure, though Picscout may have superior image-recognition algorithms, that part isn't the crowned jewels. Indeed, there are many companies with image-recognition algorithms, Google being one of them.

The real challenge is to build a network protocol that can communicate image information between a client and a server as quickly as possible, using as little network bandwidth as possible. Then, this mechanism needs to scale up to service huge volumes of requests from huge numbers of applications on the net. Picscout may be the first to introduce the proof-of-concept and a prototype, but the real race is on the back-end... as David pointed out.

On the surface, this would seem difficult -- and it is -- but it's hardly new. All large-scale social-network sites do this on a regular basis, from twitter to facebook to Flickr. Though cloud-computing is mature, the real barrier to entry here is the costly capital investment necessary to run such a service. There are many players in the field that already have this infrastructure. By comparison, Picscout would have a harder time ramping up to that level of computing resources than it would for a larger company to find some sort of image-recognition technology (if they don't already have one).

For now, the game is Picscout's to lose, since they're first. But "first" players often find themselves in catch-up soon thereafter. If they even moderately demonstrate viability in the concept, much larger players (such as photo-sharing sites) who have such resources already will be quick to swoop in.

Lastly, David notes in his point #3:
If buyers find it easier to find images through web search they will move away from distributor sites for search, and only use the distributor site for the final licensing.


Yes. Exactly. But that's nothing new. It's been that way since about 2002 now, a fact that I've been pounding on since that time: The vast number of licensed images are done on a peer-to-peer basis directly between buyers and photographers. Stock agencies have suffered because they've missed this point, and have since struggled in trying to figure out how to fight their way out of the paper bag.

But that struggle will end without their having to do much about it. With the combination of image-recognition and web-crawling, the emerging business model Picscout is attempting is now a Fait accompli. That is, David is correct to say that stock agencies of today will become nothing more than hosting sites and clearing houses that supply inventory to other middle-man sites (like Picscout) that do the real job of pairing buyers and sellers.

But is this really a bad thing? He says it in a way that suggests that agencies somehow preserve stock prices. Let's not forget that if ISVs and others realize there's money to be made, they don't want to under-price inventory too. If you want to preserve price stability, convert the social-networks from photo-sharing into photo-licensing businesses.

I've nothing against agencies, but their future will require them to do two things they never did before--in fact, that they avoided: rank well in search engines (so that end-users are more likely to find content in the first place), and attract as much content as possible. That is, stop being editors. Let any and all images in, and let the natural ranking abilities of search engines and social-networks be the real editors. To date, stock agencies have neither sufficient content volume or web-ranking in search results, nor do they employ social-network aspects to their sites to attract users in high volumes. (Again, their head was in the sand for too long.)

So the question is, who can do this? Answer: Photo-sharing social networks.

Back in 2008, I posted an article titled, Stock Photography, the Consumer, and the Future that forecasts this very phenomenon. Once the realization that there's lots of money to be made by creating a streamlined and automated image-licensing mechanism, the sleeping giants of the photo-sharing social networks will awaken and bulldoze over the traditional stock agencies in ways that no one would have believed.

Indeed, I wrote in January, 2008 in an article titled, Pulling the Flickr sword out of the Yahoo stone:
Flickr is one of the very few photo-asset powerhouses on the web that could monetize its content in ways that would exceed even modest expectations.
In fact, I also wrote in an article titled, The Solution to Getty's Woes that Getty should acquire Flickr for this very reason.

But times have changed considerably since then -- Getty has shrunk in size, and Yahoo! has recovered handsomely. Getty could never acquire Flickr now... but if this whole business model of using image-recognition as a vehicle for licensing images shows promise, then I wouldn't be surprised if Yahoo! starts casting devious stares towards Getty.

Hmmmm......

Labels: , , , , , , , , , ,